Geo‑Distribution
& Multi‑Region
Architecture
DYNAMODB GLOBAL · COCKROACHDB · GDPR · RPO/RTO
Active-passive failover
Active-active with replication
Read replicas in each region
Regional active-active writes
Regional data partitioning
Geo-routing + KMS per region
// Raft cluster: US-East (leader), EU-West, AP-Southeast // Quorum = 2 of 3. Client writes from US-East.Write latency = time until quorum ACK US-East → EU-West round trip: ~85ms US-East → AP-Southeast RT: ~140ms
Quorum write latency = min(EU-West RTT, AP-Southeast RTT) = ~85ms // must wait for 1 of 2 remote nodes to ACK // This means every write to your database takes 85ms minimum. // A 10ms SLA is IMPOSSIBLE with synchronous global Raft. // Solution: non-voting replicas (CockroachDB style) // US-East: 2 voting replicas (quorum = 2, all local → ~1ms write latency) // EU-West: 1 non-voting replica (async replication, ~1s lag) // AP-Southeast: 1 non-voting replica (async replication, ~1s lag) // Write latency: ~1ms (local quorum) ✓ // EU/APAC reads: served locally (slightly stale) ✓
✓ No conflict resolution needed
✓ Strong consistency always
✗ Remote users pay write latency
✗ Failover gap: 30s–minutes
✗ Risk: split-brain on failover
✓ Survives full region loss
✓ No bottleneck primary region
✗ Eventually consistent
✗ Some conflicts = silent data loss (LWW)
✗ Harder to debug
| CONFLICT STRATEGY | HOW IT WORKS | DATA LOSS? | BEST FOR |
|---|---|---|---|
| Last-Write-Wins | Highest timestamp wins on conflict | Yes — concurrent write silently lost | Sessions, caches, preferences |
| Vector Clocks | Track causality; surface conflicts to app | No — app resolves explicitly | Shopping carts (Dynamo original) |
| CRDTs | Mathematically conflict-free merge | Never — by design | Counters, sets, collaborative editing |
| Operational Transform | Transform ops against concurrent ops | Never — but complex | Real-time collaborative text (Google Docs) |
Merge: max per slot
Value: sum = 10
Value = 10 - 3 = 7
Merge: {x,y,z}
Remove is permanent
add(x, id2) → x still present
Concurrent inserts merge correctly
class GCounter:
def __init__(self, node_id, num_nodes):
self.node_id = node_id
self.counts = [0] * num_nodes # slot per node
def increment(self):
self.counts[self.node_id] += 1 # only increment OWN slot
def value(self):
return sum(self.counts)
def merge(self, other): # merge another replica
self.counts = [max(a, b)
for a, b in zip(self.counts, other.counts)]
# Simulation: 3 nodes
A = GCounter(0, 3); A.increment(); A.increment(); A.increment() # [3,0,0]
B = GCounter(1, 3); B.increment(); B.increment() # [0,2,0]
C = GCounter(2, 3); C.increment() # [0,0,1]
A.merge(B); A.merge(C) # [3,2,1] → value = 6
B.merge(A); B.merge(C) # [3,2,1] → value = 6
# All replicas converge to 6 regardless of merge order ✓
// Architecture: table replicated across N AWS regions
// Each region: accepts reads AND writes independently
// Replication: asynchronous, bidirectional, ~1s lag
// Conflict resolution: LAST-WRITE-WINS (timestamp-based)
// GOOD USES (LWW acceptable):
✓ User sessions: {user_id → session_token, last_seen}
✓ User preferences: {user_id → theme, language, notifications}
✓ Shopping cart: {user_id → cart_items} (last write wins per user)
✓ Configuration flags: {feature_flag → enabled/disabled}
// BAD USES (LWW causes data loss):
✗ Financial balances: concurrent increments → one lost
User in US adds $100, user in EU adds $50 simultaneously
One write wins → balance shows +$100 OR +$50, not +$150
✗ Inventory: concurrent decrements → overselling
✗ Leaderboard rankings: concurrent score updates
// For counters: use DynamoDB Streams + Lambda to consolidate
// Or: use a CRDT service (PN-Counter semantics)
// Or: route all writes for a given key to its "home" region
// Cost: each additional replica region ≈ 2× storage + throughput costs
// Replication lag: typically ~1s, can spike to ~5s under high load
-- Set survival goal: lose an entire region without data loss
ALTER DATABASE mydb SURVIVE REGION FAILURE;
-- REGIONAL BY ROW: each row pinned to its home region
-- EU user's rows stored in EU region → EU reads/writes at 5ms latency
ALTER TABLE users SET LOCALITY REGIONAL BY ROW AS region;
-- User in EU: crdb_region='eu-west-1' → row stored in EU → fast local access
INSERT INTO users (id, name, crdb_region) VALUES (1, 'Alice', 'eu-west-1');
-- REGIONAL TABLE: entire table anchored to one region
-- Good for: tables only accessed from one region
ALTER TABLE eu_compliance_log SET LOCALITY REGIONAL IN 'eu-west-1';
-- GLOBAL: replicated everywhere, optimized for reads everywhere
-- Writes need global consensus (slower), reads are local (fast)
-- Good for: product catalog, reference data, configuration
ALTER TABLE product_catalog SET LOCALITY GLOBAL;
-- Follower reads: slightly stale (~4.8s) but instant from any region
-- Good for: analytics, dashboards, non-critical reads
SELECT * FROM orders AS OF SYSTEM TIME follower_read_timestamp()
WHERE user_id = 123;
Non-personal: aggregated counts, anonymized analytics
JWT: {"user_id": 123, "region": "eu-west-1"}
Service: read/write user data only from claimed region
US cluster: us-east-1 only, KMS key: us-east-1
Backups: encrypted with regional KMS → stays in region
Delete: KMS.deleteKey(user_123)
Result: all stored blobs are now unreadable garbage
Good: log("user login: user_id=hash(alice) region=eu")
// Given: 99.99% availability SLA (52 min/year downtime budget)
// Incident: detect=15min, page team=5min, fix=30min → 50 min per incident
// → Budget allows ZERO incidents that go over 52 min/year
// → Need automatic failover with RTO
// Payment service (financial data):
RPO = 0 → synchronous replication (zero data loss)
RTO = 30s → hot standby, auto-promote (no manual steps)
Cost: higher write latency, expensive hot standby
// Analytics service (aggregate counts):
RPO = minutes → periodic snapshots (some loss OK)
RTO = hours → restore from snapshot (downtime OK)
Cost: cheap backup storage, no standby infra
// Product catalog (semi-static data):
RPO = seconds → async replication (small loss OK)
RTO = minutes → warm standby, semi-auto (brief outage OK)
Cost: moderate — warm standby only
- Active-active Raft cluster: US-East (leader), EU-West, AP-Southeast. Client writes from US-East. What is the minimum write latency? Show the RTT math.
- Same cluster but AP-Southeast is non-voting. Now what is the write latency? What does AP-Southeast contribute?
- Add SA-East (~90ms from US-East) as a 4th non-voting replica. Does write latency change? Does read latency for SA-East users change?
- Tokyo user on active-passive system (US-East is primary, AP-Southeast has read replica). RTT for a write? RTT for a read? What SLA is achievable for Tokyo reads vs writes?
- Implement
PNCounter(nodeId, numNodes)withincrement(),decrement(),merge(other), andvalue() - Test: Node A increments 3×, Node B increments 5× and decrements 2×. Merge A→B and B→A. Verify both show value=6.
- Test idempotency: merge A into B twice. Does the result change?
- Test commutativity: merge(A,B) == merge(B,A)?
- Why can a regular integer counter NOT be a CRDT? Which of the three properties (commutative, associative, idempotent) does it violate?
- List all personal data fields in the B6 Twitter design. Which must be GDPR-protected?
- Design regional partitioning: can an EU user's tweet be cached in a US Redis instance? Can it be in a US CDN edge node?
- User exercises right to erasure. Step-by-step deletion from: Cassandra (with TTR replicas), Redis timelines, ClickHouse analytics, Kafka topic logs.
- Design cryptographic erasure: what is the KMS key structure? Who manages keys? How long does deletion take to propagate?
- EU user's timeline includes tweets from a US user. Is including those tweets in EU storage a GDPR violation?
50M India users, 20M EU users. GDPR applies to EU. Products in both India and EU warehouses.
- Product catalog (read-heavy, non-personal): which replication strategy? Active-active LWW? Read replicas? GLOBAL locality? Justify.
- Inventory (globally shared — 1 unit in Bangalore can be bought by India or EU user): how do you prevent overselling without a global lock?
- Orders: EU orders must stay in EU (GDPR). India user buys from EU warehouse — where does the order record live?
- RPO/RTO for each service: inventory RPO=0/RTO=30s, catalog RPO=minutes/RTO=hours, orders RPO=0/RTO=60s. Design the specific replication for each.
- Right to erasure for an EU user who has orders, reviews, and browsing history. What gets deleted? What can be retained (anonymized)?
A/B testing at scale · Shadow mode deployment · Feedback loops
Real-time vs batch inference · Model versioning · Embeddings at scale