Observability Stack: Metrics, Logs, Traces, and the SRE Golden Signals
Observability is not a dashboard, a vendor, or a checkbox on a launch review. It is the practical ability to ask arbitrary questions about your running system and get useful answers without shipping new code. When an incident hits at 3 AM, the difference between a ten-minute recovery and a three-hour outage is almost always the quality of the telemetry you captured before things broke. This pillar walks you through building a real observability stack around Prometheus, Grafana, Loki, Tempo, and OpenTelemetry, applying the SRE golden signals, designing alerts that page humans only when action is required, and using error budgets to negotiate risk with product teams. Read this end to end if you want a working mental model, or jump to the section that matches your current pain.
What observability actually means
The word "observability" comes from control theory, where it describes how well internal states of a system can be inferred from external outputs. In software, the working definition is simpler: you have observability when you can debug novel problems using the data your system already emits. That last phrase matters. Monitoring answers questions you knew to ask in advance, like "is CPU above 80 percent?" Observability answers questions you did not anticipate, like "why did checkout latency spike for European customers on iOS between 14:02 and 14:07?"
The traditional three pillars are metrics, logs, and traces. Metrics are numerical time series, cheap to store and fast to query, ideal for dashboards and alerts. Logs are timestamped records of discrete events, high in cardinality and detail but expensive at scale. Traces follow a single request across service boundaries, showing where latency accumulates and where errors originate. None of these pillars is sufficient alone. A latency spike shows up in metrics, gets narrowed by traces, and gets root-caused in logs. A missing pillar means guessing.
A useful fourth signal, often called events or change events, records deployments, feature flag flips, config changes, and infrastructure mutations. Most incidents correlate with a recent change, so overlaying change events on your latency and error dashboards is the fastest root cause tool most teams never build. If you take one thing from this page, wire your CI/CD system to emit a deployment event to your metrics store on every rollout.
Observability is a socio-technical practice, not just a stack. The best Prometheus setup in the world will not help if engineers do not know how to read a flame graph, if on-call runbooks are stale, or if alerts fire so often that people mute the channel. Treat tooling and culture as equal investments. If you are new to the operational culture side, start with the SRE fundamentals guide before spending money on vendors.
The SRE golden signals
Google's Site Reliability Engineering book codified four golden signals that should be measured on every user-facing service: latency, traffic, errors, and saturation. These four are not the only useful metrics, but they are the minimum viable dashboard for a production service. If you cannot answer all four for a service in under thirty seconds, you have a gap.
Latency is how long requests take. Always measure it as a distribution, never as an average. Averages hide tail behavior, and users experience tails. Track p50, p95, p99, and p99.9 separately, and split successful requests from failed ones. A 500ms failure is not the same as a 500ms success. In Prometheus, use histograms with explicit buckets tuned to your SLO, not summaries, because histograms can be aggregated across instances.
Traffic is demand on your system, measured in the units that matter for your workload. For a web service, that is requests per second. For a queue consumer, it is messages processed per second. For a database, it is queries per second broken down by type. Traffic gives context to the other three signals. A latency spike during a 10x traffic surge is a capacity story. The same spike at normal load is a bug story.
Errors is the rate of failed requests, and the trick is defining "failed" honestly. HTTP 500s count, obviously, but so do HTTP 200s that returned the wrong content, requests that succeeded but took longer than your SLO, and background jobs that silently swallowed exceptions. Instrument your service to emit a first-class error metric with a reason label, and audit what you consider success at least once a quarter.
Saturation is how full your system is, expressed as a percentage of some hard limit. CPU, memory, disk IO, network bandwidth, database connections, thread pool capacity, and queue depth are all candidates. Saturation is a leading indicator; latency and errors are lagging. A system running at 85 percent saturation for hours will eventually tip over even if latency looks fine right now. Alert on saturation trends, not just thresholds.
For request-driven services these four are enough. For batch systems, add freshness (how old is the most recent successful run) and correctness (did the output match expected invariants). For data pipelines, add completeness and lineage. Adapt the framework; do not treat it as gospel.
The RED and USE methods
Two other frameworks complement the golden signals and are worth knowing by name. The RED method, from Tom Wilkie, focuses on services and tracks Rate, Errors, and Duration. It maps almost directly to three of the four golden signals and is the right starting point for any HTTP or RPC service. If your team can produce a RED dashboard for every service, you have solved 70 percent of the "where do I look" problem during incidents.
The USE method, from Brendan Gregg, focuses on resources and tracks Utilization, Saturation, and Errors. Where RED asks "how is this service behaving," USE asks "is this resource healthy." Apply USE to CPUs, disks, network interfaces, and any bounded resource. The two methods are complementary. RED tells you the service is slow; USE tells you the underlying disk is saturated. Together, they cover the surface area of most production incidents.
A practical rule: every service dashboard should have a RED panel at the top and a USE panel for its dependencies below. This one layout, applied consistently, eliminates most of the "which dashboard do I open" confusion during on-call.
Prometheus as the metrics backbone
Prometheus has become the default open source metrics system for a reason. Its pull-based model, dimensional data model, and PromQL query language fit cloud-native workloads well, and the exporter ecosystem covers almost every off-the-shelf component you will run. If you are starting fresh, Prometheus with a long-term storage backend like Thanos, Cortex, or Mimir is the safe default.
Prometheus scrapes HTTP endpoints exposing metrics in a simple text format. Each metric is a name plus a set of key-value labels, and each label combination produces a distinct time series. This dimensionality is powerful but dangerous. A metric like http_requests_total{method, path, status, user_id} will explode into millions of series if user_id is unbounded. The rule of thumb is to keep label cardinality under a few thousand per metric, and never label with unbounded identifiers like user IDs, request IDs, or full URLs.
PromQL takes practice. The core operators to master are rate() for counters, histogram_quantile() for latency percentiles, sum by (label) for aggregation, and topk() for finding outliers. A useful exercise: reproduce your existing dashboards in PromQL from scratch, without copying queries. You will discover which panels you actually understand. For a structured introduction, the free Prometheus learning path walks through queries, alerting rules, and recording rules with concrete examples.
Recording rules deserve special mention. If a PromQL query is expensive or reused across dashboards, precompute it with a recording rule that writes the result back as a new time series. This trades a small amount of storage for large query speedups and keeps dashboards responsive during incidents, which is exactly when responsiveness matters most. A common pattern is to precompute per-service p95 latency and error rate once per minute, then reference the recorded metric in every dashboard and alert.
Retention and storage are the operational hard part. Prometheus stores data locally by default, and local storage is fine for two weeks of high-resolution data. For anything longer, ship to object storage via Thanos or Mimir. Do not try to run a single giant Prometheus instance for a whole fleet. Federation and sharding by service or team scale better than vertical growth.
Grafana for dashboards and exploration
Grafana is the visualization layer that ties everything together. It queries Prometheus, Loki, Tempo, and dozens of other sources, and it renders dashboards, ad hoc explorations, and alert states in one interface. The mistake most teams make is treating Grafana as a place to accumulate dashboards forever. A dashboard graveyard is worse than no dashboards, because on-call engineers waste time on stale panels during incidents.
Adopt a dashboard hierarchy. At the top, a small number of executive or fleet-wide dashboards show SLO compliance and error budget burn across all services. In the middle, one service dashboard per team follows a consistent template: RED metrics on top, saturation for key dependencies in the middle, deep-dive links at the bottom. At the bottom, deep-dive dashboards drill into specific subsystems and are linked from the service dashboard, not discovered by search.
Templating with variables is where Grafana pays for itself. A single dashboard template with $service, $environment, and $region variables replaces hundreds of copy-pasted dashboards and stays consistent as your fleet grows. Combined with dashboards-as-code stored in Git alongside your infrastructure as code, you get version history, code review, and reproducibility for your observability surface.
Annotations are the underused feature that transforms Grafana from a viewer into an investigator. Configure your CI/CD pipeline to post deployment annotations to every relevant dashboard. Configure your feature flag system to post flag change annotations. Configure your incident management tool to post incident start and end annotations. When a latency spike appears, the annotation overlaid on the graph tells you the cause in one glance instead of thirty minutes of correlation.
Structured logging with Loki
Logs are the pillar most teams overinvest in and underuse. Unstructured text logs are cheap to write and expensive to search. The remediation is structured logging: every log line is a JSON object with a consistent schema including timestamp, level, service, trace ID, span ID, and event-specific fields. Once your logs are structured, you can query them like a database instead of grepping like a caveman.
Loki, from Grafana Labs, takes a different approach than Elasticsearch-based stacks. Instead of indexing log content, Loki indexes only labels and stores the raw log stream in object storage. This makes it dramatically cheaper at scale, at the cost of slower content searches. For most operational use cases, where you know the service and time range and want to find a specific error, Loki is faster and cheaper than an inverted index. For full-text search across a year of logs, an Elasticsearch-based system still wins.
The label strategy for Loki mirrors Prometheus. Keep cardinality low. Use labels for service, environment, level, and pod name. Do not use labels for trace ID, user ID, or request ID; those go in the log line itself and are found via LogQL filters. LogQL, Loki's query language, supports both label selectors and content grep, and can compute metrics from log streams via rate() and similar functions. This last capability, metrics from logs, is useful when you have not yet instrumented a metric but the information exists in log output.
Log volume is the operational constraint. A busy service can emit gigabytes of logs per hour, and most of it is useless during incidents. Adopt log levels rigorously. Debug logs go to stdout in dev and are filtered out in production. Info logs record business events and state transitions. Warn logs record recoverable problems worth investigating. Error logs record failures with enough context to reproduce. If every request produces ten info lines, you have a sampling problem, not a search problem.
Distributed tracing with Tempo and OpenTelemetry
Traces are the pillar that pays off most in complex, distributed systems and least in monoliths. A trace follows a single request as it hops through services, capturing spans for each unit of work. Each span records start time, duration, service, operation, and attributes. A trace visualizer renders this as a flame graph, showing where time went and where errors originated. For a request that touches ten microservices, a trace can pinpoint the bottleneck in seconds. No amount of log grepping matches this.
OpenTelemetry is the vendor-neutral standard for generating traces, metrics, and logs. It has effectively won the instrumentation war. Its SDKs cover every major language, its collector normalizes and routes telemetry to any backend, and its semantic conventions standardize attribute names across ecosystems. If you are starting new instrumentation today, use OpenTelemetry. If you have existing OpenTracing, OpenCensus, or vendor-specific instrumentation, plan a migration. The lock-in cost of vendor SDKs is real, and OTel eliminates it.
Tempo, also from Grafana Labs, is the storage backend that pairs naturally with Loki and Prometheus. Like Loki, it minimizes indexing and stores raw traces in object storage, keyed by trace ID. You find a trace by ID, usually via a link from a log line or a metric exemplar, and Tempo returns the full span tree. Combined with the exemplars feature in Prometheus, which attaches sample trace IDs to histogram buckets, you can click from a latency graph to the exact slow trace that caused the spike.
The hard part of tracing is not the backend, it is achieving useful coverage. A trace is only as good as its weakest span. If one service in the chain is not instrumented, the trace has a hole and the flame graph lies. Adopt tracing service by service, prioritize the ones on hot paths, and require new services to emit traces from day one. Sampling is another lever. Head-based sampling at one percent is common and cheap, but it misses rare errors. Tail-based sampling, where the collector decides after seeing the full trace whether to keep it, catches errors and slow requests while dropping normal ones, at the cost of more collector complexity.
Context propagation is where tracing quietly breaks. HTTP requests must carry trace headers, and every intermediary, load balancer, service mesh, message queue, must preserve them. W3C Trace Context is the current standard, and any modern OTel SDK handles it, but legacy systems and homegrown RPC layers often drop headers. Audit propagation with a synthetic request that logs the trace ID at every hop.
OpenTelemetry as the unification layer
The strategic value of OpenTelemetry goes beyond tracing. The OTel Collector is a pipeline for all telemetry, capable of receiving from any source, transforming with processors, and exporting to any backend. This means you can run one agent on every host that handles metrics, logs, and traces, batch and compress them, add resource attributes like cluster and region, and ship to Prometheus, Loki, Tempo, or a commercial vendor, all from one config.
A common architecture: applications emit OTLP (OpenTelemetry Protocol) to a local collector sidecar or DaemonSet. The local collector adds resource attributes, samples traces, and forwards to a regional gateway collector. The gateway collector applies tail-based sampling, enforces rate limits, and fans out to storage backends. This two-tier design isolates application performance from backend availability. If Tempo is down, the collector buffers traces locally and retries. Applications never know.
OpenTelemetry semantic conventions are the underrated feature. They define standard attribute names for HTTP requests, database queries, messaging systems, and cloud resources. When every service in your fleet uses http.status_code and http.route consistently, cross-service dashboards and alerts become trivial. When one team uses status and another uses http_status, you spend engineering time on translation instead of investigation. Enforce conventions in code review or via collector processors that rename attributes at ingest.
Vendor-neutral instrumentation is also a career-portable skill. The specific dashboards you build for one employer do not transfer. The ability to instrument a Go service with OTel, design a sampling strategy, and debug a broken trace pipeline transfers everywhere. If you are building your DevOps career, prioritize OTel over vendor-specific SDKs. The DevOps engineer program covers OpenTelemetry instrumentation, Prometheus, and Grafana alongside the Kubernetes and CI/CD skills that make observability actionable.
Service Level Objectives and error budgets
Metrics without SLOs are just numbers. A Service Level Objective is a target for a Service Level Indicator, measured over a rolling window. "99.9 percent of homepage requests complete in under 300ms over the last 28 days" is an SLO. The SLI is the underlying measurement; the SLO is the target; the SLA, if you have one, is the contractual commitment with penalties.
The error budget is the inverse of the SLO. If your SLO is 99.9 percent, your error budget is 0.1 percent, which over 28 days is about 40 minutes of downtime or equivalent latency violations. The budget is a currency. You can spend it on risky deploys, chaos experiments, or ambitious refactors. When the budget is exhausted, you stop shipping features and focus on reliability until the budget recovers. This is the mechanism that turns reliability from a vague aspiration into an operational discipline.
The mechanics of error budget policy matter more than the numbers. Who has authority to freeze deploys? What counts as an SLO violation, and how quickly does the alert fire? Are budgets per service or per user journey? Get these answers in writing, socialize them with product and engineering leadership, and revisit them quarterly. An error budget policy that no one enforces is worse than no policy, because it burns team credibility.
Multi-window, multi-burn-rate alerts are the modern way to alert on SLOs. A single threshold alert either fires too late (after significant budget burn) or too often (on transient blips). The multi-burn-rate approach fires a page when the last hour and the last five minutes both show high burn rates, indicating a real ongoing problem, and fires a ticket for slow burns that need investigation but not immediate action. Google's SRE workbook has the canonical formulation; adopt it verbatim rather than inventing your own.
Alerting strategy: from noise to signal
The single biggest observability failure mode is alert fatigue. When on-call engineers get paged twenty times a night and eighteen are noise, they stop reading pages. The two real incidents get missed. Every alert must satisfy three criteria: it is caused by a real user-visible problem, it is actionable by the person paged, and it requires human action within minutes. If any criterion fails, it is not a page.
Split your alerts into three tiers. Pages wake people up and demand immediate response; they should number in the single digits per week per service. Tickets file into a queue for next-business-day handling; they cover slow SLO burns, capacity trends, and non-urgent errors. Dashboard signals show state without notifying anyone; they cover context that helps during incidents but does not require action. Most teams put too many alerts in the page tier and end up numb.
Symptom-based alerting beats cause-based alerting. Alert on "checkout success rate below 99 percent" rather than "database connection pool above 80 percent." The symptom alert catches every cause, including ones you did not anticipate. The cause alert misses novel failure modes and fires spuriously when the cause does not actually degrade users. Cause-based alerts belong in the ticket tier, not the page tier.
Runbooks are inseparable from alerts. Every page must link to a runbook with three sections: how to verify the alert is real (not a monitoring glitch), how to mitigate user impact quickly (rollback, failover, feature flag), and how to escalate. Runbooks are living documents; update them after every incident. If an alert fires and its runbook is wrong or missing, that is a bug to fix before the next on-call rotation.
Deduplication and grouping matter at scale. A single failing dependency can trigger fifty alerts from fifty services. Your alerting system, whether Alertmanager, PagerDuty, or another, must group related alerts into a single incident with a single owner. Otherwise the on-call is fighting the notification system instead of the outage.
Kubernetes-native observability patterns
Kubernetes changes observability because workloads are ephemeral, densely packed, and dynamically scheduled. A pod that existed for four minutes and produced the error you are debugging is gone by the time you look. This shifts weight from host-based tools to cluster-wide telemetry pipelines that capture data before pods disappear.
The standard Kubernetes stack is Prometheus with the kube-prometheus-stack Helm chart, which bundles Prometheus, Alertmanager, Grafana, node-exporter, kube-state-metrics, and preconfigured dashboards for cluster health. Deploy this on day one, before you deploy your first application workload. The Kubernetes operations guide covers cluster setup and the operational patterns that make the observability stack itself reliable.
Application logs from Kubernetes pods should be collected via a DaemonSet running a lightweight agent like Promtail, Fluent Bit, or the OTel Collector. The agent tails container log files, enriches with pod and namespace labels from the Kubernetes API, and ships to Loki or your log backend. Do not have applications log directly to a network destination; the DaemonSet pattern survives network blips and pod restarts.
Kubernetes-specific signals to alert on include pod restart rate, pending pod count, node NotReady status, PersistentVolumeClaim usage, and control plane latency. kube-state-metrics exposes all of these, and the kube-prometheus-stack ships default alerts for most. Review those defaults; some are too noisy for production and need threshold tuning.
Service meshes like Istio or Linkerd add another layer, emitting L7 metrics and traces for every service-to-service call without application code changes. This is the fastest way to get RED metrics across a large fleet, and it is especially valuable during a migration to OpenTelemetry when application instrumentation is incomplete. The tradeoff is operational complexity; a service mesh is another distributed system to run.
Cost control and cardinality management
Observability costs scale with cardinality and volume, and both grow silently. A well-meaning engineer adds a user_id label to a metric, and next month your Prometheus bill doubles. A new service logs every request body, and Loki storage triples. Cost surprises are almost always cardinality surprises. Instrument cost monitoring on your observability stack itself; the meta-observability layer is what keeps the bill sane.
Set explicit cardinality budgets per service. A service should not emit more than, say, 10,000 unique time series without review. Enforce with Prometheus's sample_limit scrape config and with Grafana alerts on active series growth. When a team wants to exceed the budget, they should justify it, not sneak it in.
Log volume follows the same logic. Set per-service log volume budgets, alert when they are exceeded, and require teams to sample or drop verbose logs rather than expand storage. Structured logging helps here: you can drop debug-level logs at the collector and keep info and above, saving 80 percent of volume with no loss of signal.
Trace sampling is the cost lever with the most subtle tradeoffs. One percent head-based sampling is cheap but misses rare errors. Tail-based sampling that keeps all errors and slow traces plus one percent of normals is more expensive to run but preserves the traces that matter. For most teams, tail-based sampling at the gateway collector is the right default once trace volume becomes material.
Retention tiers save money without losing analytical capability. Keep high-resolution data for 14 to 30 days for incident response. Downsample to lower resolution for 90 days to a year for trend analysis. Archive to cold storage for compliance if required. Thanos, Mimir, and commercial vendors all support this pattern natively. Do not pay hot-storage prices for data no one queries.
Incident response and the observability stack
The purpose of observability is to reduce Mean Time To Recovery. Every design choice should be evaluated against a single question: does this help an on-call engineer at 3 AM find and fix the problem faster? Vanity dashboards and clever queries do not; consistent templates, clear runbooks, and fast queries do.
The first minute of an incident is the highest-leverage. In that minute, the on-call needs to answer: is this real, is it me or a dependency, and how bad is it? Design your top-level dashboards to answer these three questions in one page load. SLO burn rate, error rate by service, and dependency health should be visible without scrolling or clicking.
Correlation is the middle-game skill. Once you know the service is degraded, you need to correlate the degradation with a cause. Deployment annotations, feature flag changes, dependency errors, and traffic anomalies are all candidates. The exemplars pattern, where Prometheus attaches trace IDs to metric samples, lets you jump from "p99 spiked" to "here is a specific slow trace" in one click. This one workflow, well implemented, cuts investigation time in half.
Postmortems close the loop. Every significant incident should produce a written postmortem covering timeline, impact, root cause, contributing factors, and action items. The observability angle in every postmortem is the same question: what data did we wish we had, and how do we capture it next time? The action items from this question, over many incidents, are how your observability stack matures. Skip this step and you rebuild the same gaps forever. For the broader operational culture that makes postmortems useful, see the SRE fundamentals guide.
Chaos engineering is postmortems in advance. Deliberately break parts of your system in controlled experiments and observe how your telemetry responds. If a broken dependency does not show up on your dashboards within a minute, that is a monitoring gap to fix before real users find it. Game days, tabletop exercises, and automated chaos tooling all serve this purpose. The observability stack is only as good as its response to real failures, and chaos experiments are how you test it without waiting for outages.
Building the stack: a reference architecture
Here is a concrete reference architecture for a small to medium production environment, roughly 50 to 500 services. Adapt to your scale.
| Layer | Component | Purpose |
|---|---|---|
| Instrumentation | OpenTelemetry SDKs | Emit metrics, logs, traces from apps |
| Collection | OTel Collector (DaemonSet) | Local aggregation, resource enrichment |
| Collection | OTel Collector (gateway) | Sampling, routing, rate limiting |
| Metrics storage | Prometheus + Thanos or Mimir | Short and long-term metrics |
| Log storage | Loki | Structured log storage on object storage |
| Trace storage | Tempo | Trace storage on object storage |
| Visualization | Grafana | Dashboards, exploration, alerting UI |
| Alerting | Alertmanager, PagerDuty | Route pages and tickets |
| Change events | CI/CD webhooks to Prometheus | Deployment annotations |
The step-by-step rollout order, if you are starting fresh, is:
- Deploy Prometheus and Grafana with kube-prometheus-stack. Get cluster and node metrics working.
- Instrument one canary service with OpenTelemetry metrics. Build a RED dashboard for it.
- Define an SLO for that service. Wire up multi-burn-rate alerts.
- Deploy Loki and Promtail. Get structured logs from the canary service.
- Deploy Tempo and add tracing to the canary. Wire exemplars from metrics to traces.
- Roll out the pattern to the next five services. Codify the template.
- Set cardinality and volume budgets. Instrument the observability stack itself.
- Integrate CI/CD deployment annotations and feature flag events.
- Write runbooks for every page-tier alert. Do a game day to test them.
- Repeat the rollout to remaining services, prioritizing by user-facing risk.
This order front-loads value. By step 3 you have a working SLO for one service. By step 6 you have a repeatable template. Teams that try to boil the ocean by instrumenting fifty services in parallel usually stall; teams that ship one great example and copy it succeed.
Tie the whole stack to your delivery pipeline. Your CI/CD pipeline should emit deployment events, and your GitOps workflow should version dashboards and alerting rules alongside application code. Observability configuration drift is as dangerous as infrastructure drift, and the same tools that solve it for infra solve it here.
Career and learning path
Observability skills are portable across employers, cloud providers, and language ecosystems in a way that many DevOps skills are not. A senior engineer who can design an SLO, tune a multi-burn-rate alert, debug a broken OTel pipeline, and lead an incident review is valuable at any company running production software. These skills are also in growing demand as more organizations adopt SRE practices.
Building this skill set requires deliberate practice, not just reading. Run your own stack, ideally on a homelab or a small cloud account. Instrument a real application, break it, debug it, and iterate on the telemetry. Contribute to an open source project's observability layer; the OpenTelemetry, Prometheus, and Grafana communities all welcome contributors. Take on the on-call rotation at your current job and volunteer for postmortem facilitation.
For structured learning that combines observability with the surrounding DevOps skills, the Refonte DevOps engineer program walks through Prometheus, Grafana, OpenTelemetry, Kubernetes, and CI/CD as an integrated curriculum with mentored internship work on real systems. If you prefer to plan your own path, the DevOps hub links out to every subtopic, and the DevOps certifications guide covers which credentials actually signal skill to hiring managers. For a broader view of where the field is heading, the DevOps trends and career guide covers adjacent skills like platform engineering and FinOps.
Common anti-patterns to avoid
Dashboard-driven observability. Building dashboards without SLOs produces pretty graphs that no one uses during incidents. Start with the question "what does user pain look like," then build the SLO, then build the dashboard.
Vendor lock-in through SDKs. Using a commercial vendor's proprietary instrumentation SDK ties you to that vendor forever. Use OpenTelemetry SDKs and send data to the vendor via OTLP. You can switch backends without touching application code.
Alert on everything. More alerts do not equal more safety; they equal more fatigue. Every alert must be actionable. Prune quarterly.
Ignoring cardinality. A single high-cardinality label can double your metrics bill overnight. Set budgets and enforce them.
Logs as a database. If you are querying logs to compute business KPIs, you have a data pipeline problem, not a logging problem. Emit metrics or events to a proper analytics system.
Traces without context propagation. A trace that ends at your load balancer because headers were dropped is worse than no trace, because it lies about where time went. Audit propagation end to end.
Reactive-only observability. Waiting for incidents to reveal gaps is expensive. Run game days and chaos experiments to find gaps proactively.
One dashboard per engineer. Personal dashboards proliferate and go stale. Enforce templated, team-owned dashboards that follow a consistent layout.
FAQ
How much should we spend on observability? A reasonable benchmark is 5 to 10 percent of your total infrastructure spend. Under 3 percent usually means gaps that cost more in outages than they save. Over 15 percent usually means uncontrolled cardinality or excessive log retention. Measure the ratio and treat it as a budget line.
Prometheus or a commercial vendor like Datadog? Both are valid. Prometheus with Thanos or Mimir gives you full control and lower per-GB costs at scale, but you operate it yourself. Datadog, New Relic, or Honeycomb give you turnkey capability at higher unit cost. The right choice depends on team size, expertise, and how many engineers you can dedicate to running the stack. Many teams run Prometheus for metrics and a commercial vendor for traces or logs.
Do we need distributed tracing if we have logs and metrics? If you have more than three services on a request path, yes. Traces answer "where did time go across services" in seconds; logs and metrics answer it in hours. If you have a monolith, tracing is nice to have but not urgent.
How do we handle observability across multiple clouds? Standardize on OpenTelemetry for instrumentation. Run a collector in each cloud that ships to a central backend, or run per-cloud backends federated at the query layer with Thanos or Grafana. Do not use cloud-specific instrumentation SDKs; you will regret it when you migrate or add a second cloud.
What is the difference between monitoring and observability? Monitoring watches known failure modes with predefined checks. Observability lets you investigate unknown failure modes with rich telemetry. In practice, mature teams do both: monitoring drives alerts, observability drives debugging.
How do we get engineers to actually instrument their code? Make it the default. Provide a shared library or template that emits OTel telemetry with zero code from the service author. Require SLOs and RED dashboards as launch criteria. Review telemetry quality in code review. Culture beats tooling; leadership must treat observability as non-optional.
When should we adopt SLOs? As soon as you have a service with real users. You do not need a mature observability stack to define an SLO; you need one SLI you can measure and a target. Start with one SLO per critical user journey and expand from there.
What metrics matter for a database? Query latency percentiles, query rate by type, error rate, connection pool utilization, replication lag if applicable, disk saturation, and cache hit rate. The USE method applies cleanly to databases; every bounded resource should have utilization and saturation panels.
