What Is Concurrency in a Distributed System?

Concurrency in a distributed system means that many operations are in progress at the same time on different machines, and those machines share no memory and no common clock. A single server with 200 worker threads is concurrent. A system of 50 such servers behind a load balancer is concurrent and distributed, and the second case is harder because no node can see what the others are doing at this instant. The practical problem is keeping shared data correct when two nodes change it at the same time.

Concurrency, parallelism, and distribution

Concurrency means that several tasks overlap in time. Parallelism means that several tasks execute at the same instant on separate processors. Distribution means the tasks run on separate machines connected by a network. A distributed system is always concurrent, usually parallel, and never able to rely on shared memory.

The missing shared memory is the important part. Two threads on one machine can protect a counter with a mutex, a lock that only one thread holds at a time, because both threads see the same memory. Two services on different machines can only send messages, and every message can be delayed, duplicated, or lost. Every concurrency control in a distributed system is a way to agree on order without shared memory.

What goes wrong without control

Lost updates. Two order services read a stock count of 1 for the same item. Both subtract 1 and both write 0. Two customers receive a confirmation for the last unit.

Reordered operations. A user changes an email address twice in one second. The two updates travel over different network paths and arrive at the database in reverse order. The older address wins.

Double processing. A payment worker crashes after charging a card but before recording the charge. The queue redelivers the message and a second worker charges the card again.

Deadlock. Service A holds a lock on the account table and waits for the ledger table. Service B holds the ledger table and waits for the account table. Neither can finish, and on separate machines neither can detect the cycle without help.

Four ways to handle concurrency

Pessimistic locking with leases. A distributed lock is a record in a shared store, such as Redis or ZooKeeper, that one node holds at a time. The lock has a lease, an expiry of perhaps 10 seconds, so a crashed holder does not block everyone forever. Use it when conflicts are frequent and the protected work is short.

Optimistic concurrency control. Each record has a version number. A writer reads version 7, computes the new value, and writes only if the version is still 7. The database rejects the write if another node has moved the record to version 8, and the writer reads again and retries. Use it when conflicts are rare, which is the common case for user-facing records.

UPDATE inventory SET quantity = quantity - 1, version = version + 1 WHERE sku = 'A1' AND version = 7 AND quantity > 0;

A result of zero rows updated means another node changed the record first. The application reads the record again and decides whether to retry or to tell the customer the item is gone.

Transactions and isolation. A transaction groups several reads and writes so that they succeed or fail as one unit. Isolation levels state how much concurrent transactions may see of each other. Serializable isolation behaves as if transactions ran one after another and costs the most throughput. Read committed allows more concurrency and permits some anomalies. Across machines, a two-phase commit coordinates the decision, at the price of one extra network round trip and a blocking coordinator.

Single writer per partition. The simplest control is to remove the conflict. Assign every key to one owner. In Kafka, one partition is consumed by exactly one consumer in a group, so events for one order are processed in order by one process. In a sharded database, one shard owns each account. Concurrency still exists across partitions, but never inside one.

Idempotency: the rule that makes retries safe

Retries are unavoidable in a distributed system because a timeout does not say whether the operation happened. An operation is idempotent when doing it twice has the same effect as doing it once. Setting a balance to 40 is idempotent. Subtracting 10 is not. The common technique is an idempotency key: the client sends a unique id with the request, the server stores it with the result, and a repeat of the same id returns the stored result without acting again. Payment APIs such as Stripe work this way.

Where this appears in cloud systems

A web tier with 50 instances handles concurrent HTTP requests, and each instance may hold hundreds of threads. A managed database such as DynamoDB exposes conditional writes, which are optimistic concurrency control under another name. A message queue such as SQS delivers a message at least once, so every consumer must be idempotent. A stream processor such as Flink processes one partition per task to keep order. The same four controls appear under different product names.

ControlConflict rate it suitsCostExample
Lock with leaseHighWaiting, lock serviceLeader election, batch job ownership
Version checkLowRetries on conflictProfile edits, inventory decrement
TransactionAny, within one databaseLower throughput, coordinationMoney transfer between two accounts
Single writerAnyCareful key designKafka partitions, sharded accounts

Key Takeaways

  • Distributed concurrency has no shared memory. Every control is a way to agree on order through messages.
  • Pick the control by conflict rate. Version numbers for rare conflicts, locks with leases for frequent ones, single writer when you can partition.
  • Every consumer must be idempotent. Networks retry, so a repeated operation must not repeat its effect.
  • Transactions are cheap inside one database and expensive across machines. Keep a transaction inside one shard where possible.

The consistency, replication, and messaging chapters of Grokking System Design Fundamentals cover these controls with diagrams.

Grokking the System Design Interview applies them to full designs such as a ticket booking system and a payment service.

For thread-level concurrency inside one process, see Grokking Multithreading and Concurrency for Coding Interviews.

TAGS
System Design Interview
System Design Fundamentals
CONTRIBUTOR
Arslan Ahmad
Arslan Ahmad
ex-FAANG engineering manager and author or Grokking series.

GET YOUR FREE

Coding Questions Catalog

Design Gurus Newsletter - Latest from our Blog
Boost your coding skills with our essential coding questions catalog.
Take a step towards a better tech career now!
Explore Answers
How to prepare for Mongodb system design interview for experienced individuals?
What to Expect in the Two Sigma System Design Interview
Two Sigma's design evaluation runs through its design-and-implementation round: you architect a small system and then build it working, plus data-platform design at research scale.
What do Microsoft coders do?
What are the categories of design patterns?
Design patterns fall into three categories: creational, structural, and behavioral. Concurrency is often added as a fourth.
What questions are asked in a design interview?
Which visa is best for a software engineer?
Related Courses
New
Grokking the AI System Design Interview course cover
Grokking the AI System Design Interview
Learn to design AI systems the way interviewers expect: classic ML products, LLM and RAG architectures, and agentic systems, all through the lens of the system design interview.
4.6
(3,192 learners)
Discounted price for Your Region

$123

Grokking the Coding Interview: Patterns for Coding Questions course cover
Grokking the Coding Interview: Patterns for Coding Questions
The 24 essential patterns behind every coding interview question. Available in Java, Python, JavaScript, C++, C#, and Go. The most comprehensive coding interview course with 543 lessons. A smarter alternative to grinding LeetCode.
4.6
Discounted price for Your Region

$197

Grokking Modern AI Fundamentals course cover
Grokking Modern AI Fundamentals
Master the fundamentals of AI today to lead the tech revolution of tomorrow.
4.1
Discounted price for Your Region

$72

Design Gurus logo
One-Stop Portal For Tech Interviews.
Copyright © 2026 Design Gurus, LLC. All rights reserved.