Module B6 β Design Twitter/X Feed
System Design Mastery Course | Track B: HLD | Week 16
π― Module Overview
Duration: 1 Week | Track: B β HLD Case Studies Prerequisites: B1βB5 (all fundamentals + URL Shortener) Goal: Twitter/X Feed is the canonical FAANG interview question for social systems. It forces you to make the single hardest trade-off in feed design: fan-out on write vs fan-out on read. Master both models and the hybrid.
1. Requirements
Functional
1. Post a tweet (text β€ 280 chars, optional media)
2. Follow / unfollow users
3. Home timeline: latest ~200 tweets from people I follow
4. User timeline: all tweets from a specific user
5. Search tweets (full-text)
6. Trending topics / hashtags
7. Likes, retweets, replies (counts)
Non-Functional (Scale)
Users: 300M total, 100M DAU
Tweets/day: 500M (5,800/sec avg, 15,000/sec peak)
Timeline reads: 28B/day β 320K reads/sec avg, 800K/sec peak
Read:Write ratio ~50:1 (timeline reads vs tweet posts)
Follows: avg 200 followers, some celebrities have 100M+
Storage/tweet: ~280 bytes text + metadata β 1 KB total
Storage/day: 500M Γ 1 KB = 500 GB/day
5yr total: ~900 TB (text only), multi-PB with media
Latency targets:
Home timeline load: p99 < 200ms
Post tweet: p99 < 500ms
Search results: p99 < 1s
2. The Core Problem: Home Timeline
The home timeline is the hard part. Given user U who follows 500 people, each of whom tweets ~3Γ/day:
Naive approach: at read time, fetch tweets from all 500 followees, merge, sort.
500 followees Γ DB query each = 500 queries per timeline load
At 320K timeline reads/sec: 160M DB queries/sec β IMPOSSIBLE
Better approach: pre-compute the timeline. But HOW?
3. Fan-Out on Write (Push Model)
When user A tweets, push the tweet ID to the inbox of every follower.
User A (200 followers) posts tweet T:
For each follower F in A.followers:
timeline:F.prepend(T.id) β Redis sorted set, score = timestamp
Timeline read for user U:
ZREVRANGE timeline:U 0 199 β O(1) read from Redis
Batch fetch tweet content: MGET tweet:id1 tweet:id2 ...
Data Flow
[POST /tweet] β Kafka topic "tweet-created"
β
β
[Fanout Worker Service]
reads user A's followers from social graph DB
for each follower F: ZADD timeline:F timestamp tweetId
β
β
[Redis timeline cache]
timeline:userId β sorted set of tweet IDs (latest 1000)
Trade-offs
β
Timeline read is instant β O(1) Redis lookup
β
Scales reads to millions/sec easily
β
Consistent experience β all followers see tweet immediately after fanout
β Write amplification: 1 tweet Γ 100M followers = 100M Redis writes
β Celebrities (Lady Gaga, 100M followers) cause massive write spikes
β Wasted writes: pushed to timelines of inactive users who never open app
β Timeline storage: 300M users Γ 1000 tweet IDs Γ 8 bytes = 2.4 TB in Redis
4. Fan-Out on Read (Pull Model)
At read time, pull tweets from all followees and merge.
Timeline read for user U (follows 500 people):
1. Fetch U's follow list: SELECT followee_id FROM follows WHERE follower_id = U
2. For each followee F: fetch recent tweets (last 200)
3. Merge all tweet streams by timestamp, take top 200
4. Return to user
Trade-offs
β
No write amplification β 1 tweet = 1 DB write
β
Works perfectly for celebrities (no 100M-follower write spike)
β
No wasted storage for inactive user timelines
β Read is expensive: 500 DB queries per timeline load
β 320K reads/sec Γ 500 queries = 160M queries/sec β impossible at scale
β Latency is high: merging 500 streams takes 100β500ms
β Hot users' tweet tables are hotspots for reads
5. Hybrid Approach (What Twitter Actually Uses)
Combine both models based on follower count:
Rule:
Normal users (< 10K followers): fan-out on WRITE
Celebrities (β₯ 10K followers): fan-out on READ (lazy inject at read time)
POST tweet by @normalUser (800 followers):
β Fanout service pushes tweet ID to 800 followers' Redis timelines immediately
β Fast fanout, manageable write cost
POST tweet by @ladygaga (100M followers):
β Tweet stored in DB + tweet cache only
β NO immediate fanout (would require 100M Redis writes)
READ timeline for user U:
1. Fetch pre-computed timeline from Redis (fan-out-on-write tweets)
2. Check follow list for celebrities U follows
3. Fetch recent tweets from each celebrity (fan-out-on-read component)
4. Merge the two sets by timestamp
5. Return top 200
This limits celebrity fanout reads to ~number of celebrities U follows (~few dozen max)
6. Data Model
Tweets Table (MySQL sharded by user_id)
CREATE TABLE tweets (
tweet_id BIGINT PRIMARY KEY, -- Snowflake ID (encodes timestamp)
user_id BIGINT NOT NULL, -- author
content VARCHAR(280),
media_ids JSON, -- array of S3 keys
reply_to_id BIGINT, -- NULL if not a reply
retweet_of BIGINT, -- NULL if not a retweet
like_count BIGINT DEFAULT 0, -- approximate, updated async
retweet_count BIGINT DEFAULT 0,
created_at TIMESTAMP,
INDEX (user_id, created_at DESC) -- user timeline query
);
Follows Table (separate social graph service)
CREATE TABLE follows (
follower_id BIGINT NOT NULL,
followee_id BIGINT NOT NULL,
created_at TIMESTAMP,
PRIMARY KEY (follower_id, followee_id),
INDEX (followee_id) -- "who follows @ladygaga?"
);
Redis Timeline Cache
Key: timeline:{userId}
Type: Sorted Set (ZSet)
Score: tweet timestamp (Unix epoch ms)
Member: tweet_id
ZADD timeline:123 1700000000 tweet_id_abc
ZREVRANGE timeline:123 0 199 β top 200 most recent tweet IDs
Max size: 1000 entries per user (ZREMRANGEBYRANK to trim)
Tweet Content Cache (Redis)
Key: tweet:{tweetId}
Type: Hash
Fields: userId, content, likeCount, retweetCount, createdAt
TTL: 24 hours
HGETALL tweet:{id} β hydrate tweet for display
7. Architecture
[Client]
β
ββ POST /tweet βββ [Tweet Service] βββ [MySQL: tweets shard]
β ββββββββββββ [Kafka: tweet-created]
β β
β [Fanout Service]
β reads social graph
β pushes to Redis timelines
β (skips celebrities)
β
ββ GET /home-timeline βββ [Timeline Service]
β ββ ZREVRANGE timeline:{userId} from Redis
β ββ inject celebrity tweets (fan-out-on-read)
β ββ merge by timestamp
β ββ MGET tweet:{id} for each (batch hydrate)
β
ββ GET /user-timeline βββ [Tweet Service]
β SELECT * FROM tweets WHERE user_id=? ORDER BY created_at DESC
β
ββ POST /follow βββ [Social Graph Service] βββ [Graph DB / MySQL: follows]
β
ββ GET /search βββ [Search Service] βββ [Elasticsearch index]
8. Media Storage
Tweets with images/videos:
Upload: Client β CDN upload endpoint β S3 (origin)
Serve: Client β CDN edge (cached) β S3 (on miss)
Tweet stores: media_ids: ["s3://tweets/2024/01/img_abc.jpg"]
CDN URL: https://pbs.twimg.com/media/img_abc.jpg
Video transcoding:
Raw upload β S3 raw bucket β Lambda/Worker β transcode to HLS (multiple bitrates)
β S3 transcoded bucket β CDN
Why CDN is essential:
Viral tweet with 100M impressions Γ 500KB image = 50 TB transferred
Without CDN: origin S3 + bandwidth bill is astronomical
With CDN: 99%+ cache hit rate for popular media
9. Search & Trending
Search (Elasticsearch)
On tweet creation β Kafka β Search indexer consumer β Elasticsearch index
Index fields: content (full-text), userId, timestamp, likeCount
Query: GET /search?q=ukraine+war
Elasticsearch: full-text match + rank by (recency + engagement)
Response: top 50 tweet IDs β hydrate from tweet cache
Trending Topics
Approach: sliding window count of hashtags
Implementation:
Kafka stream: extract hashtags from each tweet
Flink/Storm: count per hashtag in 1-hour sliding window
Top-K: maintain min-heap of top 30 hashtags
Store: Redis sorted set "trending" β ZADD trending count #hashtag
Read: ZREVRANGE trending 0 9 β top 10
Refresh rate: every 5 minutes
Geographic trending: separate sorted set per region (trending:US, trending:IN)
10. Counts (Likes, Retweets, Followers)
Problem: 500M tweets/day Γ likes/retweets = billions of write ops/day
Can't UPDATE tweets SET like_count += 1 synchronously at this rate
Solution: Counter service with async aggregation
1. User likes tweet β POST /like
2. Immediately: write to likes table (for "did I like this?" check)
3. Async: Kafka event "tweet-liked" β counter aggregation worker
4. Counter worker: batches 1000 events β UPDATE tweets SET like_count += N
Or: Redis INCR like_count:{tweetId} β periodic flush to DB
Display: serve from Redis cache (approximate) β exact count in DB
Accuracy: ~1-5 second lag. Acceptable β Twitter shows "1.2M likes" not "1,234,567"
π Tasks
Task 1 β Fan-Out Analysis
Calculate the cost of fan-out on write for these scenarios:
- 100M-follower celebrity tweets once β how many Redis writes?
- 100K users each follow this celebrity β how many extra reads at timeline load?
- Whatβs the storage difference: full fan-out-on-write vs hybrid (threshold 10K followers)?
Task 2 β Data Model Extension
Extend the schema to support:
- Twitter Lists (user-created collections of accounts)
- Tweet thread / conversation chains
- Quote tweets (retweet with comment)
- Pinned tweets (shown first on user profile)
Task 3 β Failure Scenarios
Design failure handling for:
- Fanout service crashes mid-fanout (10M of 100M followers updated, then crash)
- Redis timeline cache is wiped (cache miss for all 300M users simultaneously)
- Social graph DB is down (canβt resolve followers for fanout)
- Elasticsearch is unavailable (search requests fail)
β Task 4 β Full Design Interview Simulation
Set a 45-minute timer. On a blank sheet of paper, design Twitterβs home timeline from scratch using the 7-step framework. Include:
- Requirements + NFRs with numbers
- Capacity estimation
- Full architecture diagram
- Fan-out decision + hybrid threshold justification
- DB schema (tweets + follows)
- Redis timeline cache design
- At least 2 edge cases
- 1 scaling evolution (what changes at 10Γ current scale)
β Checklist
- Can state Twitterβs scale: 300M users, 500M tweets/day, 320K timeline reads/sec
- Explain fan-out on write: what it does, write amplification problem
- Explain fan-out on read: what it does, read fan-out problem
- Know the hybrid approach: threshold-based, celebrity inject at read time
- Can draw the complete architecture with all 6 services
- Know the Redis timeline cache design (sorted set, score = timestamp)
- Know the DB schema: tweets table + follows table + indexes
- Understand the celebrity problem and why it breaks pure fan-out-on-write
- Know how likes/counts are handled (async Kafka counter service)
- Know media storage: S3 + CDN + transcoding pipeline
- Know trending topics: sliding window, Flink, Redis sorted set
- Know search: Elasticsearch, indexing pipeline, hydration pattern
- Tasks 1β3 completed
- Task 4: Full 45-min timed simulation completed