System Design Fundamentals
What is System Design?
System Design is the process of defining the architecture, components, modules, interfaces, and data for a system to satisfy specified requirements.
💡 Why System Design Matters
System design skills are crucial for building scalable, reliable, and maintainable applications. They help you make informed decisions about architecture, technology choices, and trade-offs.
Key Concepts:
Scalability - Ability to handle increased load
Reliability - System operates correctly over time
Availability - System is operational when needed
Performance - Response time and throughput
Maintainability - Easy to modify and extend
Part 1: System Architecture Patterns
Monolithic Architecture
Benefits:
Simple to develop and deploy
Easy to test and debug
No network latency between components
ACID transactions across all data
Drawbacks:
Difficult to scale individual components
Single point of failure
Technology lock-in
Deployment affects entire system
Monolithic Architecture Structure
🏗️ Monolithic Architecture 📱 CLIENT LAYER Web Client Browser-based UI HTML/CSS/JS User interface Client-side logic Mobile Client Native Apps iOS/Android Mobile interface Native functionality 🖥️ MONOLITHIC APPLICATION Presentation Layer User Interface HTTP handlers Request routing Response formatting Business Logic Layer Core Logic Domain logic Business rules Validation Data Access Layer Database Operations ORM/SQL Query logic Data mapping 💾 DATA STORAGE Single Database Centralized Data All tables Shared schema ACID transactions Single point File Storage Static Assets Images Documents Media files 🔄 CHARACTERISTICS Single Deployment All-or-Nothing One codebase One build One deploy Tight Coupling Dependent Components Shared memory Direct calls Strong dependencies Technology Stack One Language/Framework Limited choices Team alignment Ecosystem lock-in
Part 2: Scaling Concepts
Vertical vs Horizontal Scaling
Vertical Scaling:
Add more power to existing server
Simpler architecture
Limited by hardware constraints
Higher cost per unit of performance
Horizontal Scaling:
Add more servers to the system
Better fault tolerance
More complex architecture
Better cost efficiency at scale
Scaling Request Flow Comparison
Understanding how each scaling approach handles incoming requests is crucial for architectural decisions.
📈 VERTICAL SCALING REQUEST FLOW Client 1 Request Client 2 Request Client 3 Request Powerful Server High Resources 32 CPU cores 256GB RAM Processes all requests Single instance Single Database All Connections All traffic Single point 📊 HORIZONTAL SCALING REQUEST FLOW Client 1 Request Client 2 Request Client 3 Request Load Balancer Traffic Distribution Round-robin Health checks Request routing Server 1 Instance 4 CPU cores 16GB RAM Handles subset Server 2 Instance 4 CPU cores 16GB RAM Handles subset Server 3 Instance 4 CPU cores 16GB RAM Handles subset Database Cluster Distributed Load Primary + Replicas Load balanced
Scaling Trade-offs Comparison
⚖️ Scaling Approach Trade-offs 📈 VERTICAL SCALING TRADE-OFFS Pros ✅ Simple to implement ✅ No code changes ✅ Easy to debug Cons ❌ Hardware limits ❌ Single point of failure ❌ Downtime for upgrades 📊 HORIZONTAL SCALING TRADE-OFFS Pros ✅ Unlimited scale potential ✅ Cost effective ✅ Fault tolerant Cons ❌ Complex architecture ❌ Distributed challenges ❌ Network latency 🎯 USE CASES Vertical Best For 🖥️ Small to Medium Apps Low traffic Predictable load Simple architecture Startup/MVP Legacy systems Horizontal Best For 🌐 Large Scale Apps High traffic Unpredictable load Cloud-native Microservices High availability needs 💰 COST ANALYSIS Small Scale 1-10K users Vertical: $1,000/mo Horizontal: $800/mo Winner: Tie Medium Scale 100K users Vertical: $10,000/mo Horizontal: $3,000/mo Winner: Horizontal Large Scale 1M+ users Vertical: $100,000/mo Horizontal: $20,000/mo Winner: Horizontal
Scaling Decision Flowchart
Limited Flexible < 10K users > 100K users > 1M users 10K-1M users Need to Scale? Assess System Budget Available? Expected Traffic? High Traffic? Choose Vertical Scaling 📈 Scale Up Simple upgrade Quick implementation Single server Best for low-medium traffic Choose Horizontal Scaling 📊 Scale Out Add servers Load balancer Best for high traffic Cloud-native ready Hybrid Approach ⚖️ Best of Both Vertical for database Horizontal for app servers Optimized cost Flexible scaling Check Scalability Scale Database 💾 Database Layer Read replicas Write sharding Caching layer CDN for static data Scale Application 🖥️ App Servers Load balancer Server instances Auto-scaling Container orchestration Monitor & Iterate 📊 Performance Metrics QPS monitoring Latency tracking Error rates Cost analysis
Scaling Architecture Comparison
⚖️ Vertical vs Horizontal Scaling 📈 VERTICAL SCALING - SCALE UP Single Powerful Server 🖥️ All-in-One Increase CPU cores Increase RAM Increase storage Increase bandwidth Upgrade hardware Database Single Instance More storage More memory Single point of failure Application Single Process More threads More memory Limited scalability Characteristics 💰 Expensive ⚠️ Hardware limits 🔒 Single point of failure 📊 HORIZONTAL SCALING - SCALE OUT Load Balancer ⚖️ Traffic Distribution Round-robin Weighted distribution Health checks Request routing Server 1 🖥️ Instance 1 CPU/RAM resources Independent processes Can serve traffic Server 2 🖥️ Instance 2 CPU/RAM resources Independent processes Can serve traffic Server N 🖥️ Instance N CPU/RAM resources Independent processes Can serve traffic Database Cluster 💾 Distributed Storage Primary + Replicas Data sharding High availability Characteristics 💰 Cost effective ♾️ Unlimited scale 🛡️ Fault tolerant 📊 SCALABILITY METRICS Cost per Unit 💰 Resource Efficiency Vertical: $$$$ Horizontal: $$ Cost/performance ratio Scaling Limits 🎯 Maximum Scale Vertical: Hardware max Horizontal: Software max Cloud elasticity Fault Tolerance 🛡️ Reliability Vertical: SPOF Horizontal: Redundancy Resilience
Part 3: Key Metrics
Performance Metrics
Latency Percentiles:
P50 (Median): 50% of requests
P90: 90% of requests
P95: 95% of requests
P99: 99% of requests
P99.9: 99.9% of requests
The Nines of Availability:
99% = 3.65 days downtime/year
99.9% = 8.76 hours downtime/year
99.99% = 52.56 minutes downtime/year
99.999% = 5.26 minutes downtime/year
Performance Metrics Visualization
📊 System Performance Metrics ⏱️ LATENCY PERCENTILES P50 - Median 50% of requests Typical response time User experience baseline Most common latency P90 - 90th Percentile 90% of requests Good performance target Acceptable user experience Service level metric P99 - 99th Percentile 99% of requests Exceptional cases Outlier handling System resilience P99.9 - 99.9th Percentile 99.9% of requests Edge cases Extreme scenarios Critical systems 🎯 AVAILABILITY NINES 99% Availability Uptime 99% uptime 87.6 hrs/year downtime Basic reliability 3.65 days/year downtime 99.9% Availability Uptime 99.9% uptime 8.76 hrs/year downtime Good reliability ~7 hours/year downtime 99.99% Availability Uptime 99.99% uptime 52.56 min/year downtime High reliability ~1 hour/year downtime 99.999% Availability Uptime 99.999% uptime 5.26 min/year downtime Mission-critical ~5 min/year downtime 📈 THROUGHPUT METRICS QPS - Queries Per Second Request Rate System capacity Load handling Scalability planning TPS - Transactions Per Second Transaction Rate Database capacity Transaction throughput ACID operations Bandwidth Data Transfer Rate Network capacity Data movement CDN utilization ⚠️ ERROR METRICS Error Rate Failure Percentage 4xx errors 5xx errors System health SLA - Service Level Agreement Uptime Guarantee Business commitment Penalty thresholds Customer expectations
Latency vs Throughput Relationship
Understanding how latency and throughput affect each other is critical for system performance optimization.
⚖️ LATENCY VS THROUGHPUT TRADE-OFF 🎯 OPTIMIZATION STRATEGIES 📊 IMPACT FACTORS Low Latency ⚡ Fast Response < 50ms response Real-time systems User experience Fast connections High Throughput 📈 Many Requests/sec 10K+ QPS Batch processing High capacity Parallel processing Balanced Performance ⚖️ Optimal Acceptable latency Good throughput Cost effective Scalable Network Latency 🌐 Bottleneck Round-trip time Distance Protocol overhead Packet loss Processing Speed 🖥️ CPU/RAM Algorithm complexity Cache efficiency I/O operations Resource limits System Capacity 📦 Resource Limits Server capacity Connection pools Queue depth Bandwidth Caching Strategy 💾 Reduce Latency In-memory cache CDN distribution Local caching Redis/Memcache Load Balancing ⚖️ Distribute Load Multiple servers Traffic routing Health checks Auto-scaling Database Optimization 💾 Query Tuning Indexes Read replicas Connection pooling Query optimization
Availability vs Reliability Relationship
Understanding the difference between availability and reliability helps design systems that meet business requirements.
🏛️ Availability vs Reliability 🟢 AVAILABILITY Uptime Percentage System Operational Time online / Total time Measured in 9s User-facing metric Service reachability 99% = 3.65 days/year Basic Systems Internal tools Development systems Non-critical apps 99.9% = 8.76 hrs/year Standard Systems Business applications Work hours uptime Scheduled maintenance 99.99% = 52 min/year Mission Critical Financial systems Healthcare systems E-commerce platforms 🔒 RELIABILITY Correctness Over Time System Functionality Output correctness Data integrity Fault tolerance Error handling Error Rate Failures/Bugs 4xx client errors 5xx server errors Data corruption Incorrect results Data Integrity Consistency ACID properties Transaction safety Backup/recovery Data validation ⚖️ RELATIONSHIP & TRADE-OFFS HIGH Availability LOW Reliability ⚠️ Trade-off System stays up But returns errors Eventually consistent Final consistency HIGH Reliability LOW Availability ⚠️ Trade-off Correct outputs But frequent downtime Strong consistency Maintenance windows Ideal System Both High ✅ Target High uptime Low error rate Strong guarantees Complex to achieve
SLA/SLO/SLI Hierarchy and Relationships
Understanding the relationship between SLA, SLO, and SLI is crucial for defining and measuring system quality.
📋 SLA/SLO/SLI Relationship 📜 SERVICE LEVEL AGREEMENT (SLA) Contract with Users Business Commitment Legal agreement Penalty clauses Business targets Customer-facing SLA Examples Guarantees 99.9% uptime < 200ms p95 latency < 0.1% error rate Credits if violated 🎯 SERVICE LEVEL OBJECTIVE (SLO) Internal Goal Engineering Target Measured internally Tighter than SLA Buffer for SLA Usually 4x-10x margin SLO Examples Internal Targets 99.95% uptime < 150ms p95 latency < 0.05% error rate SRE best practices 📊 SERVICE LEVEL INDICATOR (SLI) Measured Metric Raw Data Points Actual performance Quantitative metric Rolling window Historical data Latency SLI Response Time p50: 50ms p95: 150ms p99: 500ms Measured continuously Error SLI Failure Rate 4xx: 0.5% 5xx: 0.01% Timeout: 0.1% Measured continuously Availability SLI Uptime Successful requests Health check passes System operational 99.97% measured 🔄 FLOW HIERARCHY Monitoring 📊 Collect Data Metrics collection Alerting Dashboards Real-time tracking SLI Aggregation 📈 Calculate Time window Percentiles Error rates Availability % SLO Comparison 🎯 Evaluate SLI vs SLO Gap analysis Risk assessment Improvement needed SLA Buffer 🛡️ Protection SLO > SLA Safety margin Error budget Burn rate
System Health Metric Flow
Understanding how all metrics flow together to indicate overall system health.
📊 SYSTEM HEALTH METRICS FLOW Incoming Request 📥 User Traffic HTTP request API call User interaction ⏱️ Performance Metrics Response Time ⏱️ Latency P50: 50ms P90: 150ms P99: 500ms Time to first byte Throughput 📈 Capacity QPS: 10,000 TPS: 5,000 Bandwidth usage RPS ✅ Quality Metrics Success Rate ✅ Reliability 99.95% success 0.05% errors Correctness Data integrity Error Rate ❌ Failures 4xx: 0.3% 5xx: 0.02% Timeouts Failures 🟢 Availability Metrics Uptime 🟢 Operational 99.99% online Health checks pass Active monitoring Service health Health Status 🟢 System Health Green: Healthy Yellow: Degraded Red: Down Alert status ⚖️ SLO Evaluation Calculate SLI 📊 Indicators Aggregate metrics Time windows Percentiles Error rates SLO Check 🎯 Target Met? SLI >= SLO? Error budget? Within limits? Alert threshold ✅ Healthy All SLOs Met ⚠️ Degraded Some SLOs Missed 🚨 Alert SLO Violation
Part 4: Design Principles
Core Design Principles
Separation of Concerns:
Each component has a single responsibility
Clear boundaries between different concerns
Easier to understand and maintain
Abstraction:
Hide implementation details
Provide clean interfaces
Reduce complexity for users
Modularity:
Break system into independent modules
Each module can be developed separately
Easier to test and debug
Loose Coupling:
Minimize dependencies between components
Changes in one component don't affect others
More flexible and maintainable system
Layered Architecture Pattern
Understanding how to properly separate concerns through layered architecture is fundamental to system design.
📐 LAYERED ARCHITECTURE PATTERN Client Layer 👤 User Interface Web browsers Mobile apps Thick clients Presentation only Presentation Layer 🎨 API/Views HTTP handlers Request routing Input validation Response formatting View rendering Business Logic Layer 💼 Domain Logic Business rules Core logic Validation Orchestration Process coordination Data Access Layer 💾 Persistence Database queries ORM operations Data mapping CRUD operations Transaction management Database 🗄️ Storage Layer Tables Indexes Constraints Data persistence ACID compliance External Services 🌐 Third-Party Payment gateways Email services APIs Message queues External integrations 📊 Layer Communication Rules Strict Layering ⚠️ Rules Layer can only call layer below No skipping layers Upward communication via callbacks Downward through function calls Layer Benefits ✅ Advantages Clear separation Easy to understand Independent testing Technology swap Parallel development Layer Concerns 🎯 Responsibilities Each layer has single concern Thin controllers Fat models Data layer isolation Business logic centralization
Dependency Injection and Inversion
Understanding dependency management through dependency injection and inversion of control patterns.
🔌 Dependency Injection & Inversion ❌ TIGHT COUPLING - HARD DEPENDENCIES Main Class ❌ Tight Coupling Direct instantiation Hard-coded dependencies Difficult to test Cannot swap implementations Tightly bound Hard Dependency ❌ Concrete Class MySQL database File system Specific service Implementation locked Rigid coupling ✅ DEPENDENCY INJECTION - LOOSE COUPLING Main Class ✅ Loose Coupling Dependencies injected Interface-based Easy to test Can swap implementations Flexible design Interface ✅ Abstraction IDatabase interface IFileService interface IService interface Contract definition Implementation abstract MySQL Implementation 💾 Concrete Implements IDatabase Can be swapped Easy to replace Testable PostgreSQL Implementation 💾 Concrete Implements IDatabase Can be swapped Easy to replace Testable 🎯 DEPENDENCY INVERSION PRINCIPLE High-Level Module 📊 Business Logic Depends on abstractions Interface-based Domain logic Not tied to low-level Low-Level Module 🔧 Implementation Implements abstractions Concrete details Infrastructure code Database, files, network Abstraction Layer 🔗 Interface Decouples layers Contract definition Inversion point Both depend on this 📊 BENEFITS ✅ Testability Easy Mocking Inject test doubles Isolated testing Unit test friendly Fast tests ✅ Flexibility Easy Swapping Different implementations Configuration-based Pluggable architecture Runtime selection ✅ Maintainability Easy Changes Change one layer Independent updates Clear boundaries Reduced coupling
Single Responsibility Principle
Understanding how to properly separate responsibilities to create maintainable and flexible systems.
🎯 Single Responsibility Principle ❌ VIOLATION - MULTIPLE RESPONSIBILITIES Class with Multiple Jobs ❌ Bad Design User management Email sending File handling Validation Too many reasons to change Responsibility 1 Users Create user Delete user Update user Query user Responsibility 2 Email Send email Email templates Email validation Email delivery Responsibility 3 Files Upload files Delete files File validation File storage ✅ PROPER SEPARATION - SINGLE RESPONSIBILITY User Service ✅ Single Responsibility Manage users User CRUD User validation User queries ONLY user concerns Email Service ✅ Single Responsibility Send emails Email templates Email validation Delivery logic ONLY email concerns File Service ✅ Single Responsibility Upload files Delete files File validation File storage ONLY file concerns 🔗 SERVICE COORDINATION Orchestration Layer 🎼 Coordinates Uses multiple services Coordinates work Does NOT do the work Delegates to services High-level flow Creates User 📝 Flow 1. Validate user 2. Save to database 📊 BENEFITS OF SEPARATION ✅ Easy Testing Test in Isolation Mock dependencies Fast tests Clear boundaries Unit test friendly ✅ Easy Modification Change One Thing Modify one service No side effects Clear impact Safe changes ✅ Reusability Use Anywhere Service can be reused Independent deployment Clear interface Modular design
Design Principles Architecture
This diagram visualizes the four fundamental design principles and how they create maintainable, scalable systems.
🏗️ Core Design Principles Visualization 🔀 SEPARATION OF CONCERNS Separation of Concerns 📋 Each layer has ONE responsibility Presentation Layer 🎨 Display Concern UI rendering User interaction View logic only Business Logic Layer 💼 Logic Concern Business rules Validation Core processing only Data Access Layer 💾 Data Concern Persistence Queries Storage only 🎯 ABSTRACTION Abstraction 🔒 Hide Implementation Details High-Level Interface 🌟 Simple API Clean methods Easy to use No internals exposed Low-Level Implementation 🔧 Complex Details Database drivers Network protocols Hidden from users 🧩 MODULARITY Modularity 📦 Independent Modules Module A 🔐 Authentication Self-contained Independent Reusable Module B 💳 Payments Self-contained Independent Reusable Module C 📦 Inventory Self-contained Independent Reusable 🔗 LOOSE COUPLING Loose Coupling 🔌 Minimal Dependencies Service A ✅ Loosely Coupled Uses interfaces No direct deps Easy to change Shared Interface 🔗 Contract Defines behavior Both depend on this Decouples services Service B ✅ Loosely Coupled Uses interfaces No direct deps Easy to change LOOSE_COUPLING
Part 5: Back-of-Envelope Calculations
Essential Numbers Every Engineer Should Know
Key Insights:
Memory is 100x faster than SSD
SSD is 100x faster than hard disk
Network round-trip is VERY slow
Avoid disk seeks in hot paths!
QPS Calculation Example
System Design Implications:
Design for peak QPS, not average
Read-heavy system (25:1 ratio)
Need aggressive caching strategy
Database read replicas required
CDN for static content delivery
Storage Calculation Example
Storage Strategy Implications:
Media dominates storage (99.925% vs 0.075% for text)
Need different strategies for text vs media
Text: Standard database with indexing
Media: Object storage (S3) with CDN distribution
Cost considerations: Media storage is the major cost factor
QPS Calculation Flow
Step-by-step process to calculate Queries Per Second (QPS) and determine the right architecture for your system.
< 5K QPS 5K-50K QPS > 50K QPS 🎯 Start: Estimate QPS System Sizing Goal 📝 Step 1: Gather Assumptions Input Data Total users: 10M Daily active users: 20% Requests per user: 10/day Peak traffic multiple: 10x 📊 Step 2: Calculate Daily Requests Formula 10M users × 20% DAU × 10 requests/day 📊 Step 3: Calculate Average QPS Formula 20M requests / 24 hrs / 3600 sec 📊 Step 4: Calculate Peak QPS Formula 231 QPS × 10x peak multiplier = 2,310 QPS peak 📈 Step 5: Assess Scale What architecture needed? Based on peak QPS ✅ Architecture Decision 🖥️ Single Server For < 5K QPS Monolithic app Simple architecture Single database Basic caching Low cost ✅ Architecture Decision ⚖️ Load Balanced For 5K-50K QPS Multiple app servers Read replicas Aggressive caching Database sharding Medium cost ✅ Architecture Decision 🌐 Distributed System For > 50K QPS Microservices Caching layers CDN for static content Multiple datacenters Auto-scaling 🚀 Implementation Phase Build system with chosen architecture
Storage Calculation Flow
Step-by-step process to calculate storage requirements and choose the right storage strategy for your system.
🎯 Start: Estimate Storage System Capacity Goal 📝 Step 1: Gather Assumptions Input Data Total users: 100M Average records: 1K/user Record size: 1KB Growth rate: 5%/month Retention: 5 years 📊 Step 2: Calculate Base Storage Formula 100M users × 1K records × 1KB per record 📄 Step 3: Calculate Text Storage Formula 100TB × 0.075% = 75GB text data User profiles Metadata Structured data Database storage 🎬 Step 4: Calculate Media Storage Formula 100TB × 99.925% = 99.925TB media Images Videos User files Object storage 📈 Step 5: Project Growth Formula 5% monthly × 60 months Compound growth 📦 Step 6: Choose Storage Type What technology for each? 💾 Text Storage Decision Database System 75GB structured data PostgreSQL/MySQL Indexes for speed ACID compliance Relational queries ☁️ Media Storage Decision Object Storage 99.925TB media files Amazon S3 Glacier for archives CDN integration Scalable storage 📊 Final Capacity Plan 5-Year Projection Text: 75GB in DB Media: 1,875TB in S3 Total: ~1,875TB Monthly growth: 5% 💰 Cost Analysis Budget Planning Storage: $0.023/GB/month Transfer: $0.09/GB Retrieval: $0.01/GB Backup: $0.004/GB/month 📊 Total: ~$43K/month
Bandwidth Calculation Flow
Step-by-step process to calculate network bandwidth requirements and optimize data transfer costs.
🎯 Start: Estimate Bandwidth Network Capacity Goal 📝 Step 1: Gather Assumptions Input Data Peak QPS: 10,000 Response size: 10KB Upload size: 100KB Read/Write ratio: 25:1 Media requests: 30% 📊 Step 2: Calculate Response Bandwidth Formula 10,000 QPS × 10KB = 100MB/s outgoing Text responses JSON data HTML pages 📊 Step 3: Calculate Upload Bandwidth Formula 400 writes/s × 100KB = 40MB/s incoming User uploads POST requests Media files 🎬 Step 4: Calculate Media Bandwidth Formula 3,000 media QPS × 200KB = 600MB/s outgoing Images Videos Static assets 📊 Step 5: Calculate Total Bandwidth Formula Outgoing: 100 + 600 = 700MB/s Incoming: 40MB/s 🌐 Step 6: Optimize Distribution How to deliver efficiently? 📥 Incoming Traffic Upload Bandwidth 40MB/s = 320Mbps User uploads API requests Form submissions Direct to origin 📤 Outgoing Traffic Response Bandwidth 700MB/s = 5.6Gbps Text: 100MB/s Media: 600MB/s Needs optimization ☁️ CDN Optimization Geographic Distribution 600MB/s via CDN - 85% 100MB/s origin - 15% Edge caching Regional distribution 💰 Reduces costs 80% 📊 Final Bandwidth Plan Network Requirements Incoming: 320Mbps Outgoing origin: 800Mbps Outgoing CDN: 4.8Gbps Total capacity: 5.92Gbps 💰 Cost: ~$5K/month
System Design Calculations
📊 System Design Calculations ⏱️ READ/WRITE RATIOS Read-Heavy System 📈 High Read Ratio 25 reads : 1 write Aggressive caching Read replicas CDN for static content Memcache/Redis Write-Heavy System 📝 High Write Ratio More writes than reads Write-optimized DB Append-only logs Time-series databases Write amplification Balanced System ⚖️ Equal Load Similar read/write Balanced architecture Standard database Normal caching 💾 STORAGE ARCHITECTURE Text Storage 📄 Structured Data Database: 10GB Fast access ACID transactions Indexes required Low cost Media Storage 🎬 Large Files Object storage: 99.9% CDN distribution S3-like storage Lazy loading High cost Storage Split 📊 99.925% vs 0.075% Media: 99.925% Text: 0.075% Separate strategies Different technologies 🚀 PERFORMANCE NUMBERS Memory Access ⚡ Very Fast 100ns latency 10GB/s throughput Random access Volatile SSD Storage 💾 Fast 100μs latency 1GB/s throughput Sequential better Persistent Disk Storage 🐌 Slow 10ms latency 100MB/s throughput Sequential needed Cheap Network 🌐 Very Slow 1-100ms RTT Bandwidth limited Packet loss possible Network latency 📈 CAPACITY PLANNING QPS Calculation 📊 Request Planning Peak QPS estimation Read/write split Replica planning Caching strategy Storage Calculation 💾 Capacity Planning Data growth rate Retention period Media vs text Cost estimation Bandwidth Calculation 🌐 Network Planning Data transfer rate Geographic distribution CDN utilization Cost optimization
Part 6: CAP Theorem
Understanding CAP Theorem
CAP Theorem states that in a distributed computer system, you can only guarantee two out of the following three properties:
Consistency : All nodes see the same data simultaneously
Availability : The system remains operational
Partition Tolerance : The system continues despite network failures
💡 CAP Theorem Insight : Network partitions are inevitable in distributed systems, so you must choose between consistency and availability during a partition. This fundamental trade-off shapes all distributed system design decisions.
CAP Theorem Venn Diagram
🎯 CAP Theorem Triangle CONSISTENCY Consistency 🔒 Same Data Everywhere All nodes see same data Strong consistency ACID transactions Linearizability Atomic updates Examples Single-node systems ACID databases Strong consistency models Synchronous replication AVAILABILITY Availability 🟢 System Always Up High uptime Read/write always succeed 99.9%+ uptime Operational guarantee No denial of service Examples CDNs Caching layers Elastic systems Auto-scaling PARTITION TOLERANCE Partition Tolerance 🌐 Network Failure Handling Continues despite partitions Tolerance of net splits Message loss tolerance Partial system operation Split-brain handling Examples Distributed systems Multi-datacenter Geographically distributed Cloud systems CA SYSTEM - CONSISTENCY + AVAILABILITY CA System ✅ Consistency ✅ Availability ❌ No Partition Tolerance Traditional RDBMS Single datacenter No network partitions assumed Example: MySQL in single DC CP SYSTEM - CONSISTENCY + PARTITION TOLERANCE CP System ✅ Consistency ❌ Limited Availability ✅ Partition Tolerance Strong consistency during partitions May reject writes Examples: MongoDB, HBase Zookeeper, etcd AP SYSTEM - AVAILABILITY + PARTITION TOLERANCE AP System ❌ Eventual Consistency ✅ High Availability ✅ Partition Tolerance Accepts writes during partition Resolves conflicts later Examples: DynamoDB, Cassandra CouchDB ❌ IMPOSSIBLE - ALL THREE Impossible ❌ Cannot have All Three Cannot guarantee all 3 Fundamental limitation Must choose trade-offs Design decision required
CAP Theorem Trade-offs in Depth
⚖️ CAP Theorem Trade-off Analysis 🎯 CA - CONSISTENCY + AVAILABILITY CA Design Single Datacenter No network partitions assumed Traditional database Perfect consistency Always available Single point of failure Pros ✅ Strong consistency ✅ High availability ✅ Simple architecture Cons ❌ No partition tolerance ❌ Single datacenter only ❌ Scaling limitations CA Examples 🔍 Real World Traditional MySQL PostgreSQL single node Single-instance Redis Local file systems In-memory databases 🎯 CP - CONSISTENCY + PARTITION TOLERANCE CP Design Partition Tolerant Consistency Strong consistency always Rejects requests during partition Leader election Quorum-based writes Split-brain prevention Pros ✅ Guaranteed consistency ✅ Handles network failures ✅ No data conflicts Cons ❌ Reduced availability ❌ Request rejection possible ❌ Complex leader election CP Examples 🔍 Real World MongoDB with replicas HBase Zookeeper etcd Strongly-consistent systems 🎯 AP - AVAILABILITY + PARTITION TOLERANCE AP Design High Availability Distributed Always accepts requests Eventual consistency Conflict resolution Merge strategies High uptime Pros ✅ High availability ✅ Fault tolerant ✅ Geographic distribution Cons ❌ Eventual consistency ❌ Conflict resolution needed ❌ Data staleness possible AP Examples 🔍 Real World DynamoDB Cassandra CouchDB Amazon S3 DNS system
CAP Theorem Decision Flow
Yes, Single DC No, Distributed Yes No, Availability Priority Design Distributed System 📊 CAP Decision Needed Can You Avoid Network Partitions? Choose CA System 🎯 Consistency + Availability Single datacenter No network failures Traditional database Perfect consistency Examples: MySQL, PostgreSQL Partition Inevitable ⚠️ Must Choose Trade-off Consistency Critical? Choose CP System 🎯 Consistency + Partition Tolerance Strong consistency Reject requests during partition Leader election Examples: MongoDB, HBase Zookeeper, etcd Choose AP System 🎯 Availability + Partition Tolerance Always accept writes Eventual consistency High availability Examples: DynamoDB, Cassandra CouchDB, Riak Consistency Requirements Strong Consistency 📌 Critical Financial systems Health records Legal data ACID required Eventual Consistency 📌 Acceptable Social media User preferences Cached data Non-critical
CAP Theorem Real-World Examples
🌐 Real-World CAP Theorem Examples 💾 DATABASE SYSTEMS MySQL Single Node 🔵 CA System Consistency Availability No partition tolerance MongoDB Replica Set 🔴 CP System Strong consistency Partition tolerant Availability traded Cassandra Cluster 🟡 AP System High availability Partition tolerant Eventual consistency ☁️ CLOUD SERVICES Amazon S3 Object Storage 🟡 AP System High availability Global distribution Eventual consistency DynamoDB NoSQL Database 🟡 AP System Always writes High uptime Conflict resolution RDS PostgreSQL Single Region 🔵 CA System ACID compliance High uptime Single datacenter 🔄 DISTRIBUTED SYSTEMS Kafka Message Broker 🔴 CP System Strong ordering Partition tolerance Leader election Redis Cluster Caching Layer 🟡 AP System High availability Replication Eventual sync etcd Config Store 🔴 CP System Raft consensus Strong consistency Leader election 📊 SYSTEM DESIGN PATTERN CA Pattern Single DC Internal tools CRM systems Legacy apps Monoliths CP Pattern Strong Consistency Payment systems Order processing Inventory management Financial data AP Pattern High Availability Social feeds User sessions Cache layers Content delivery
Summary
Key Takeaways:
System Design is about making informed architectural decisions
Scalability can be vertical (scale up) or horizontal (scale out)
Performance metrics include latency, throughput, and availability
Design principles guide architectural decisions
CAP Theorem defines fundamental trade-offs in distributed systems
Back-of-envelope calculations help estimate system requirements
Next Steps:
Practice with real-world examples
Learn about specific technologies and patterns
Understand trade-offs between different approaches
Build systems and measure their performance
Apply CAP theorem to design decisions
Remember: System design is an iterative process. Start simple, measure, and evolve based on real requirements and constraints.