M17 — Observability & Hardening
| Pillar | Question Answered | Data Type | Tools |
|---|---|---|---|
| Logs | "What happened, exactly?" | Discrete events with context | ELK, Loki, Fluentd |
| Metrics | "How fast / how many / how full?" | Aggregated numeric time-series | Prometheus, Grafana, Datadog |
| Traces | "Why was this request slow?" | Causal chains across services | Jaeger, Tempo, Zipkin, Honeycomb |
Logs are the cockpit voice recorder — full narrative of what was said. Metrics are the flight data recorder — altitude, speed, attitude plotted over time. Traces are the air traffic control replay — the full path of the aircraft from departure to destination. An accident investigation uses all three.
| Monitoring | Observability | |
|---|---|---|
| Approach | Predefined thresholds and dashboards for known failure modes | Ability to ask arbitrary questions about system behavior |
| Limits | Only catches failures you anticipated and built alerts for | Enables debugging novel, unknown failure modes |
| Data | Aggregated metrics, simple health checks | Logs + metrics + traces with high cardinality |
| Tooling | Nagios, simple dashboards | OpenTelemetry, Honeycomb, Grafana + Loki + Tempo |
| Module | Topic | Key Concepts |
|---|---|---|
| M17 (this) | Observability & Hardening | Logs, metrics, traces, alerting, SLO, security, rate limiting |
| M18 | Performance Engineering | Profiling, flame graphs, memory analysis, benchmark methodology |
/* Bad: unstructured text — unsearchable */ [2026-03-27 14:23:01] INFO: Order ord-9821 placed by user u-44 for $49.99
| Field | Format | Purpose |
|---|---|---|
ts | ISO-8601 with ms | Timeline reconstruction |
level | DEBUG/INFO/WARN/ERROR/FATAL | Log level filtering |
service | service name | Multi-service log aggregation |
trace_id | hex string | Correlate with traces |
span_id | hex string | Correlate with specific span |
msg | human-readable | Event description |
| Level | Use When |
|---|---|
DEBUG | Verbose detail for local dev only. Never in production — log volume explosion. |
INFO | Normal business events: request received, order placed, job started. |
WARN | Degraded but recoverable: retry succeeded, cache miss, approaching limit. |
ERROR | Unexpected failure requiring attention: DB timeout, invalid state, downstream error. |
FATAL | Unrecoverable — process will exit after logging. |
| Never Log | Why | Alternative |
|---|---|---|
| Passwords, API keys, tokens | Log aggregators, retention, and breach exposure | Log presence/absence, not value |
| Full credit card numbers | PCI-DSS violation | Log last 4 digits only |
| PII (emails, SSN, full name) | GDPR/CCPA violation | Log user_id (opaque reference) |
| Full request/response bodies | Volume, PII risk | Log status codes and latency only |
Health probe hits (/health/*) | Thousands/min of noise | Filter at log aggregator |
","level":"ERROR","msg":"admin escalation can forge log entries. Escape or use parameterized logging./* Database call log */ {“service”:“order-svc”,“trace_id”:“4bf92f3577b34da6”,“msg”:“db query”,“duration_ms”:12}
/* In Loki/Elasticsearch: search by trace_id to see full request timeline */ {trace_id=“4bf92f3577b34da6”} | json | sort by ts
| Type | Properties | Example | PromQL Usage |
|---|---|---|---|
| Counter | Monotonically increasing, never decreases, resets to 0 on restart | http_requests_total{method="GET",status="200"} | rate(http_requests_total[5m]) → requests/sec |
| Gauge | Point-in-time value, can go up or down | active_connections, memory_usage_bytes, queue_depth | Direct: active_connections > 1000 |
| Histogram | Samples bucketed by value; provides _count, _sum, _bucket | request_duration_seconds{le="0.1"} | histogram_quantile(0.99, rate(request_duration_seconds_bucket[5m])) |
| Summary | Pre-computed quantiles on client side (less flexible) | request_duration_seconds{quantile="0.99"} | Direct quantile access; can't re-aggregate across instances |
- Rate: requests per second — is traffic normal?
rate(http_requests_total[5m]) - Errors: error rate — are users experiencing failures?
rate(http_requests_total{status=~"5.."}[5m]) / rate(http_requests_total[5m]) - Duration: latency percentiles — is it slow?
histogram_quantile(0.99, rate(request_duration_seconds_bucket[5m]))
- Utilization: % time resource is busy
rate(process_cpu_seconds_total[1m]) * 100 - Saturation: extra work queued (can't keep up)
node_load1 / count(node_cpu_seconds_total{mode="idle"}) by (instance) - Errors: error count or rate of resource
node_disk_io_time_weighted_seconds_total
/metrics endpoint on a configurable interval (typically 15s). Your service exposes metrics in the Prometheus text format:
{user_id="..."} creates one time-series per user — millions of time-series destroy Prometheus. Use low-cardinality labels: method, status, endpoint (grouped). Never use user IDs, trace IDs, or UUIDs as labels.| What You Want | PromQL Query |
|---|---|
| Request rate (req/sec) | rate(http_requests_total[5m]) |
| Error rate % | rate(http_requests_total{status=~"5.."}[5m]) / rate(http_requests_total[5m]) * 100 |
| p99 latency | histogram_quantile(0.99, sum(rate(request_duration_seconds_bucket[5m])) by (le)) |
| CPU usage % | 100 - (avg by(instance)(rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100) |
| Memory used | node_memory_MemTotal_bytes - node_memory_MemAvailable_bytes |
| Kafka consumer lag | kafka_consumer_group_lag{topic="orders",partition="0"} |
| DB connection pool saturation | pg_stat_activity_count / pg_settings_max_connections |
A span represents a single unit of work within a trace (one service call, one DB query). Each span has:
trace_id— shared across all spans in the same requestspan_id— unique to this operationparent_span_id— the span that triggered this one (null for root span)- Start time + duration
- Attributes (key-value context)
- Status (OK / ERROR)
Flame chart: wider = longer. The DB SELECT at 12ms is the hot spot.
traceparent from incoming request, create a child span (using the span_id as parent_span_id), set the new span's span_id, and propagate the updated traceparent in outgoing calls.
- SDK: available in C, Go, Java, Python, etc.
- OTLP: OpenTelemetry Protocol — exports to any backend
- OTel Collector: receives OTLP, processes (batch, sample), exports to Jaeger/Tempo/Datadog
- Auto-instrumentation: inject tracing without changing application code (Java agent, eBPF)
- Manual instrumentation: create custom spans for business logic
| Strategy | How | Trade-off |
|---|---|---|
| Head-based: Always-on | Keep 100% of traces | Very expensive at scale |
| Head-based: Probability | Keep N% (e.g. 1%) | Simple, but misses rare errors |
| Head-based: Rate-limit | Keep up to N traces/sec | Bounded cost; may drop bursts |
| Tail-based: Error sampling | Buffer all traces; keep if trace has an error span | Catches errors; high memory buffer |
| Tail-based: Latency threshold | Keep if trace duration > P99 threshold | Catches slowness; complex to implement |
-
alert: HighP99Latency expr: | histogram_quantile(0.99, sum(rate(request_duration_seconds_bucket{service=“order-svc”}[5m])) by (le)) > 2.0 for: 5m labels: severity: warning
-
alert: ServiceDown expr: up{service=“order-svc”} == 0 for: 1m labels: severity: critical
- SLI (Service Level Indicator): a specific measurable metric — e.g., availability = (successful requests) / (total requests)
- SLO (Service Level Objective): internal target — e.g., availability ≥ 99.9% over 30 days. Engineering commits to this.
- SLA (Service Level Agreement): external contract with customers — stricter legal/financial penalties. SLO should be tighter than SLA as a safety buffer.
- Error Budget: SLO headroom — 99.9% SLO = 0.1% budget = 43.8 min/month. Track burn rate.
- Fast burn: consuming 14× normal rate → will exhaust budget in 1h → page immediately
- Slow burn: consuming 3× normal rate → will exhaust in ~5 days → ticket
- Alert on symptoms, not causes: Alert on high error rate (symptom users feel), not on "CPU at 80%" (cause — may not impact users)
- Every alert needs a runbook: include
runbook_urlannotation. On-call engineers should never face an alert without documented response steps. - Avoid alert fatigue: if the same alert fires weekly and engineers silence it, it's not actionable. Remove or fix it.
- Use
forduration: require condition to be sustained before paging (avoids flapping on 1-second spikes) - Group related alerts: AlertManager can group 50 firing alerts into one notification — prevents notification flood during outages
| Vulnerability | C/Backend Example | Prevention |
|---|---|---|
| SQL Injection | "SELECT * FROM users WHERE id=" + user_id | Parameterized queries only: PQexecParams(conn, "SELECT... WHERE id=$1", 1, NULL, params, ...) |
| Command Injection | system("ls " + user_input) | Never use system() with user input. Use execv() with argument array. |
| SSRF | Service fetches URL from user request body; attacker uses http://169.254.169.254/ (AWS metadata) | Allowlist of permitted outbound domains; block RFC-1918 and link-local addresses |
| Broken Access Control | User A can read User B's orders by changing order_id in request | Check authorization on every resource: WHERE id=$1 AND user_id=$2 |
| Security Misconfiguration | Debug endpoints enabled in prod, default credentials, verbose error messages | Separate prod config; disable /debug endpoints; return generic errors |
| Insecure Deserialization | Deserializing untrusted binary input (msgpack, protobuf from user) | Validate schema; set max sizes; reject unknown fields |
| Cryptographic Failures | MD5 for passwords, ECB mode, hardcoded keys | bcrypt/Argon2 for passwords; AES-256-GCM for encryption; libsodium |
/* Returns 1 if IP is RFC-1918 / link-local / loopback (block these) */ static int is_private_ip(const char ip) { struct in_addr addr; if (!inet_pton(AF_INET, ip, &addr)) return 0; uint32_t n = ntohl(addr.s_addr); return (n >> 24 == 10) / 10.0.0.0/8 / || (n >> 20 == (172 << 4) + 1) / 172.16.0.0/12 / || (n >> 16 == (192 << 8) + 168) / 192.168.0.0/16 / || (n >> 24 == 127) / 127.0.0.0/8 / || (n >> 16 == (169 << 8) + 254); / 169.254.0.0/16 */ }
int safe_fetch_url(const char url) { / 1. Parse hostname from URL (simplified) */ char hostname[256]; sscanf(url, “https://%255[^/]”, hostname);
/* 2. Allowlist check: only permitted domains */ const char *allowed[] = { “api.stripe.com”, “hooks.slack.com”, NULL }; int permitted = 0; for (int i = 0; allowed[i]; i++) if (strcmp(hostname, allowed[i]) == 0) { permitted = 1; break; }
if (!permitted) { fprintf(stderr, “SSRF blocked: %s not in allowlist\n”, hostname); return -1; }
/* 3. DNS resolution + IP check */ struct addrinfo *res; getaddrinfo(hostname, NULL, NULL, &res); char ip[INET6_ADDRSTRLEN]; inet_ntop(AF_INET, &((struct sockaddr_in *)res->ai_addr)->sin_addr, ip, sizeof(ip)); freeaddrinfo(res);
if (is_private_ip(ip)) { fprintf(stderr, “SSRF blocked: %s resolved to private IP %s\n”, hostname, ip); return -1; }
/* 4. Make the actual HTTP request / return 0; / proceed */ }
| Approach | Security | Details |
|---|---|---|
| Hardcoded in source | ❌ Never | Committed to git, all developers see it, forever in history |
| Environment variables | ⚠️ Acceptable | Not in code but visible in process env, logs, crash dumps — use only with Kubernetes Secrets |
| Kubernetes Secrets | ✅ Good | Base64 in etcd (encrypt etcd at rest); mounted as files or env; access controlled by RBAC |
| HashiCorp Vault | ✅✅ Best | Dynamic secrets (generated on request, auto-expire), audit log, lease renewal, fine-grained access control |
| AWS Secrets Manager / GCP Secret Manager | ✅✅ Best | Managed service equivalent; auto-rotation; IAM-controlled access |
int validate_order_request(const order_request_t req) { / Size bounds */ if (strlen(req->order_id) != 36) return -1;
/* UUID format: 8-4-4-4-12 hex chars with dashes */ if (!is_valid_uuid(req->order_id)) return -1;
/* Business rule bounds */ if (req->amount <= 0.0 || req->amount > 100000.0) return -1; if (req->item_count < 1 || req->item_count > 100) return -1;
return 0; /* valid / } / Use allowlist validation, not denylist: know what’s valid and reject everything else, rather than trying to enumerate all invalid inputs */
- API Gateway: global rate limiting per client IP or API key — protects all services
- Per-service: self-defense against gateway bypass or internal traffic spikes
Key property: allows bursts up to capacity while maintaining an average rate of r req/sec.
t=0: [████████████████████] 20 tokens → burst of 20 requests: OK t=0.1 [░░░░░░░░░░░░░░░░░░░░] 0 tokens → request: REJECT (429) t=0.5 [█████░░░░░░░░░░░░░░░] 5 tokens → 5 requests: OK t=1.0 [██████████░░░░░░░░░░] 10 tokens → 10 requests: OK Steady state: 10 req/sec sustained (burst allowed up to 20)
Formula:
count = prev_window_count × overlap_fraction + curr_window_count
long prev_count = redis_get_counter(client_key, prev_window); long curr_count = redis_incr_counter(client_key, curr_window, window_sec);
return (long)(prev_count * overlap + curr_count); }
| Algorithm | Burst Handling | Memory | Accuracy | Best For |
|---|---|---|---|---|
| Fixed Window Counter | Double burst at boundary (end+start of adjacent windows) | O(1) | Low (boundary problem) | Simple low-traffic systems |
| Sliding Window Log | Exact | O(requests in window) | Exact | Low-volume, exact limits needed |
| Sliding Window Counter | Approximate (±0.1%) | O(1) | High | Most production APIs |
| Token Bucket | Allows bursts up to capacity | O(1) | High | APIs tolerating short bursts |
| Leaky Bucket | Smooths all bursts, strict output rate | O(1) | High | Traffic shaping (network) |
429 with Retry-After so well-behaved clients back off correctly. Silently dropping causes clients to retry faster (thundering herd)./* HTTP request counter: label dimensions = {method, status} */ typedef struct { _Atomic(long) get_2xx, get_4xx, get_5xx; _Atomic(long) post_2xx, post_4xx, post_5xx; } http_counters_t;
/* Latency histogram: fixed buckets in seconds */ #define BUCKET_COUNT 7 static const double BUCKETS[BUCKET_COUNT] = { 0.005, 0.01, 0.025, 0.05, 0.1, 0.5, 1.0 };
typedef struct { _Atomic(long) bucket[BUCKET_COUNT]; _Atomic(long) count; _Atomic(double) sum; /* note: atomic double ops may need mutex on older C */ } latency_histogram_t;
/* Global metrics state */ static http_counters_t g_http = {0}; static latency_histogram_t g_lat = {0}; static _Atomic(long) g_active_conn = 0;
static inline void record_request(const char method, int status, double dur_s) { / Increment counter by method+status */ if (strcmp(method, “GET”) == 0) { if (status < 300) atomic_fetch_add(&g_http.get_2xx, 1); else if (status < 500) atomic_fetch_add(&g_http.get_4xx, 1); else atomic_fetch_add(&g_http.get_5xx, 1); }
/* Update histogram / for (int i = 0; i < BUCKET_COUNT; i++) if (dur_s <= BUCKETS[i]) atomic_fetch_add(&g_lat.bucket[i], 1); atomic_fetch_add(&g_lat.count, 1); / sum: use mutex for double precision (simplified: use long microseconds) / atomic_fetch_add(&g_lat.count, 0); / placeholder */ }
/* Render /metrics response body into buf */ static inline int render_metrics(char *buf, size_t sz) { int n = 0; n += snprintf(buf + n, sz - n, ”# HELP http_requests_total Total HTTP requests\n” ”# TYPE http_requests_total counter\n” “http_requests_total{method=“GET”,status=“2xx”} %ld\n” “http_requests_total{method=“GET”,status=“4xx”} %ld\n” “http_requests_total{method=“GET”,status=“5xx”} %ld\n”, atomic_load(&g_http.get_2xx), atomic_load(&g_http.get_4xx), atomic_load(&g_http.get_5xx));
n += snprintf(buf + n, sz - n,
”# HELP active_connections Current connections\n” ”# TYPE active_connections gauge\n” “active_connections %ld\n”, atomic_load(&g_active_conn));
n += snprintf(buf + n, sz - n,
”# HELP request_duration_seconds Latency histogram\n” ”# TYPE request_duration_seconds histogram\n”); for (int i = 0; i < BUCKET_COUNT; i++) n += snprintf(buf + n, sz - n, “request_duration_seconds_bucket{le=“%.3f”} %ld\n”, BUCKETS[i], atomic_load(&g_lat.bucket[i])); n += snprintf(buf + n, sz - n, “request_duration_seconds_bucket{le=“+Inf”} %ld\n” “request_duration_seconds_count %ld\n”, atomic_load(&g_lat.count), atomic_load(&g_lat.count)); return n; }
typedef struct { char trace_id[33]; /* 128-bit hex / char span_id[17]; / 64-bit hex */ } trace_ctx_t;
/* Thread-local trace context */ static _Thread_local trace_ctx_t tl_trace = {0};
static inline void log_set_trace(const char *trace_id, const char *span_id) { strncpy(tl_trace.trace_id, trace_id, 32); strncpy(tl_trace.span_id, span_id, 16); tl_trace.trace_id[32] = tl_trace.span_id[16] = ‘\0’; }
static inline const char *get_iso8601(char *buf, size_t n) { struct timespec ts; clock_gettime(CLOCK_REALTIME, &ts); struct tm *tm = gmtime(&ts.tv_sec); int len = strftime(buf, n, “%Y-%m-%dT%H:%M:%S”, tm); snprintf(buf + len, n - len, ”.%03ldZ”, ts.tv_nsec / 1000000); return buf; }
/* JSON-escape a string (handles quotes and backslashes) */ static inline void json_escape(char *dst, size_t dsz, const char *src) { size_t d = 0; for (size_t s = 0; src[s] && d + 2 < dsz; s++) { if (src[s] == ’”’ || src[s] == ’\’) dst[d++] = ’\’; dst[d++] = src[s]; } dst[d] = ‘\0’; }
#define LOG(level, msg_fmt, …) do {
char _ts[32], _msg[512], _esc[512];
get_iso8601(_ts, sizeof(_ts));
snprintf(_msg, sizeof(_msg), msg_fmt, ##VA_ARGS);
json_escape(_esc, sizeof(_esc), _msg);
fprintf(stdout,
”{“ts”:“%s”,“level”:“%s”,“service”:“order-svc”,” </span>
""trace_id”:“%s”,“span_id”:“%s”,“msg”:“%s”}\n”,
_ts, level, tl_trace.trace_id, tl_trace.span_id, _esc);
} while(0)
#define LOG_INFO(fmt, …) LOG(“INFO”, fmt, ##VA_ARGS) #define LOG_WARN(fmt, …) LOG(“WARN”, fmt, ##VA_ARGS) #define LOG_ERROR(fmt, …) LOG(“ERROR”, fmt, ##VA_ARGS)
/* Usage: / / log_set_trace(“4bf92f3577b34da6a3ce929d”, “00f067aa0ba902b7”); / / LOG_INFO(“order placed order_id=%s amount=%.2f”, order_id, amount); */
typedef struct { _Atomic(long) tokens_us; /* tokens * 1e6 (avoid float atomics) / _Atomic(long) last_refill_us; / last refill time in microseconds / long capacity_us; / max tokens * 1e6 / long rate_us; / tokens added per microsecond * 1e6 */ } token_bucket_t;
static inline long now_us(void) { struct timespec ts; clock_gettime(CLOCK_MONOTONIC, &ts); return ts.tv_sec * 1000000LL + ts.tv_nsec / 1000; }
static inline void tb_init(token_bucket_t tb, double rate_per_sec, double capacity) { tb->capacity_us = (long)(capacity * 1e6); tb->rate_us = (long)(rate_per_sec); / tokens/sec → tokens/us = rate/1e6 */ atomic_store(&tb->tokens_us, tb->capacity_us); atomic_store(&tb->last_refill_us, now_us()); }
/* Returns true if request is allowed; false if rate-limited */ static inline bool tb_allow(token_bucket_t *tb) { long now = now_us(); long last = atomic_exchange(&tb->last_refill_us, now); long elapsed_us = now - last;
/* Add tokens for elapsed time: tokens += rate * elapsed_us / long new_tokens = tb->rate_us * elapsed_us / 1000000; if (new_tokens > 0) { long current = atomic_fetch_add(&tb->tokens_us, new_tokens * 1000000LL); / Cap at capacity */ if (current + new_tokens * 1000000LL > tb->capacity_us) atomic_store(&tb->tokens_us, tb->capacity_us); }
/* Try to consume one token / long one_token = 1000000LL; long prev = atomic_fetch_sub(&tb->tokens_us, one_token); if (prev >= one_token) return true; / allowed / / Not enough tokens: restore / atomic_fetch_add(&tb->tokens_us, one_token); return false; / rate limited */ }
/* Usage: / / token_bucket_t per_client_bucket; / / tb_init(&per_client_bucket, 100.0, 200.0); // 100 req/sec, burst=200 / / if (!tb_allow(&per_client_bucket)) { respond_429(); return; } */
/metrics on port 8081 alongside /health/live.docker run -p 9090:9090 -v $(pwd)/prometheus.yml:/etc/prometheus/prometheus.yml prom/prometheus. Configure prometheus.yml to scrape localhost:8081/metrics every 15s.wrk -t4 -c100 -d30s http://localhost:8080/orders. Watch metrics accumulate at http://localhost:9090.docker run -p 3000:3000 grafana/grafana. Add Prometheus as data source. Build a RED dashboard with three panels: request rate, error rate %, p99 latency.trace_id (UUID) + span_id. If the request already has traceparent, extract the trace_id and create a child span_id.traceparent header. Service B logs with the same trace_id from the header.docker run -p 16686:16686 jaegertracing/all-in-one). Use the OTel C SDK to export spans to Jaeger and view the trace waterfall.429 Too Many Requests with Retry-After header when rate-limited.http_requests_total{status="4xx"}.wrk -t8 -c100 -d60s at 1000 req/sec (10× limit). Verify roughly 100 req/sec succeed and the rest 429. The rate should be stable over the 60s window.clang -fsanitize=thread.char query[256]; snprintf(query, sizeof(query), "SELECT * FROM orders WHERE id='%s'", user_input); PQexec(conn, query);. Try input: '; DROP TABLE orders; --. Verify it executes.PQexecParams(conn, "SELECT * FROM orders WHERE id=$1", 1, NULL, params, NULL, NULL, 0). Retry the injection — verify it returns no results (treats the entire input as a literal string).http://169.254.169.254/latest/meta-data/. Verify it returns data.safe_fetch_url() from Tab 6. Verify the metadata URL is blocked. Verify a legitimate allowlisted URL succeeds.- Explain the 3 pillars and what question each answers
- Write a structured JSON log line with all mandatory fields
- List what must never appear in logs (secrets, PII)
- Explain how trace_id links logs across services
- Distinguish counter, gauge, histogram, summary
- Write PromQL for request rate, error rate %, and p99 latency
- Apply RED method to a service and USE method to a resource
- Explain why high-cardinality labels are dangerous in Prometheus
- Write a Prometheus alerting rule with
forduration
- Explain trace, span, parent_span_id relationship
- Parse and construct a W3C traceparent header
- Choose the right sampling strategy for given traffic/budget
- Define SLI, SLO, SLA, Error Budget
- Write a burn rate alert (faster than threshold-crossing alerts)
- Explain why alerting on symptoms is better than causes
- Fix SQL injection with parameterized queries in libpq
- Block SSRF with allowlist + RFC-1918 IP check
- Explain broken access control with a concrete example
- Describe dynamic secrets (Vault) vs static env vars
- Write allowlist input validation for a struct field
- Implement token bucket: capacity, rate, burst
- Compare sliding window counter vs fixed window (boundary problem)
- Return correct 429 + Retry-After + X-RateLimit-* headers