AWS Well-Architected: The Six Pillars and How to Actually Review
The AWS Well-Architected Framework gives you a common language and a practical set of questions to evaluate your workloads against proven practices. Teams use it to uncover risk, make pragmatic improvements, and build confidence that systems are ready for real-world demand and failure. This guide shows you how to run an AWS Well-Architected Review from first invite to final remediation, and how to go deeper than a checklist. You will find detailed checklists for each pillar, examples that reflect the messy reality of mixed tech stacks, and tooling tips that keep the review fast and evidence-based. If you are looking to make an aws well architected review a regular engineering muscle, this is your field manual.
What the Well-Architected Framework is and why it matters
The AWS Well-Architected Framework is a set of design principles, architectural best practices, and review questions organized across six pillars: Operational Excellence, Security, Reliability, Performance Efficiency, Cost Optimization, and Sustainability. It helps you measure a specific workload against risks and improvement opportunities, not as a pass-fail audit, but as a guided conversation that produces an improvement plan. AWS publishes detailed pillar guidance and a standardized questionnaire to drive consistent, comparable reviews. You can apply the same approach whether your workload is a monolith on EC2, a containerized microservice platform on EKS, or a fully managed serverless application.
The immediate value is clarity. A Well-Architected Review surfaces implicit assumptions, missing guardrails, and brittle design choices that get lost in day-to-day delivery. Engineers and product leaders align around the same view of risk, with evidence captured in a way that survives handoffs. Because the questions are the same across teams, leaders can compare risk posture across portfolios, which makes budget and staffing decisions more grounded.
The framework is also vendor-native. It focuses on how to use AWS services safely and efficiently, not only on abstract design goals. For example, rather than simply saying you should encrypt data at rest, the Security pillar asks if you use KMS keys with appropriate key policies, rotation practices, and separation of duties. Where other frameworks are philosophy-forward, Well-Architected is operations-forward. If you want to read the official reference, start with the AWS documentation for the framework overview at https://docs.aws.amazon.com/wellarchitected/latest/framework/welcome.html.
Well-Architected is not certification prep, but the mindset and practices strongly reinforce the knowledge you use to pass and apply cloud credentials. If you want to build those skills systematically, you may also find value in the broader cloud learning paths at the cloud skills hub and the practical route maps in cloud certifications and career paths.
How a Well-Architected Review works, end to end
A Well-Architected Review is a structured workshop focused on one workload. It is part interview, part evidence collection, and part design improvement session. The output is a list of high-risk issues, recommendations, and owners for remediation. Plan for 2 to 6 hours of workshops plus some asynchronous evidence collection, depending on workload scope and team readiness.
Here is an end-to-end outline you can adapt:
1) Preparation and scoping - Define the workload. Name the business capability, boundary, primary users, data domains, and in-scope AWS accounts and regions. - Inventory the architecture. If you have a diagram, baseline it. If not, sketch the minimal flow from request ingress to data persistence, then enumerate critical dependencies. - Identify the review team. Include a product owner, a lead engineer, someone who runs operations, and a security representative. If you use IaC, invite the platform owner. - Agree on goals. Do you need to pass a customer due diligence check, lower costs, or get ready for a scale event? Goals shape depth and follow-up.
2) Set up the Well-Architected Tool - Use the AWS Well-Architected Tool to create the workload and track answers. The tool maintains pillar questions and flags risks. - Decide if you will share the workload with a partner or with other accounts for collaboration. - If you need more background, see the AWS user guide for the Well-Architected Tool at https://docs.aws.amazon.com/wellarchitected/latest/userguide/what-is-wellarchitected-tool.html.
3) Evidence collection - Gather CloudTrail, Config, VPC, and IAM data to verify claims during the workshop. - Export Cost Explorer views and tag coverage reports to support the Cost pillar. - Pull CloudWatch Metrics and alarm lists for critical components, and a list of backups or EBS snapshots for the Reliability pillar.
4) Facilitate the workshop - Move pillar by pillar, but follow natural dependencies. For example, discuss identity, then network segmentation before diving into app-level secrets. - Ask each review question, capture the answer, and request proof for anything subjective. - When a risk appears, write it down as a concrete issue, not a generality. For example, "RDS production instance uses default KMS key with no rotation" is actionable. "We should improve encryption" is not.
5) Document findings and owners - Use a shared spreadsheet or a ticket template to record each issue, its risk level, owner, and due date. - Align on prioritization rules before you assign due dates, to avoid underestimating security or over-focusing on short-term cost savings.
6) Follow up and verify improvements - Translate risks into changes, PRs, or tickets with acceptance criteria and testing steps. - Agree on a timeline to re-review, either for the entire workload or for specific unresolved risks.
If you are short on time, split the review across two sessions. In the first session, run Operational Excellence and Security to catch governance and guardrail gaps. In the second, finish Reliability, Performance Efficiency, Cost Optimization, and Sustainability. Avoid running the entire questionnaire in one long sitting without evidence. That typically leads to fast answers and slow improvements.
Choosing the workload to review first
You will get the most impact by reviewing a high-value workload that is actively maintained and has near-term change opportunities. The best candidate is not necessarily your largest environment, but the one with a mix of risk and decision velocity. If your organization has a critical application that has not been touched in months, you may find many risks but little appetite to fix them. Conversely, a new service under heavy development can incorporate improvements quickly, but may not yet have a clear operational baseline.
Use these criteria to select a first target: - Business impact. What happens if the workload becomes unavailable for two hours? For two days? - User exposure. Does it serve external customers, internal teams, or batch processes? - Change velocity. Do you deploy weekly or a few times per year? - Team readiness. Do you have an accountable product owner and an engineering lead who can act on findings? - Lifecycle. Are there imminent scale events or compliance reviews that justify immediate improvements?
You can formalize the decision with a simple matrix. Score each workload 1 to 5 on Impact, Exposure, and Velocity. Multiply and pick the top one or two.
| Workload | Impact | Exposure | Velocity | Score |
|---|---|---|---|---|
| Payments API | 5 | 5 | 4 | 100 |
| Data Lake Batch | 4 | 3 | 2 | 24 |
| Internal Wiki | 2 | 2 | 3 | 12 |
This scoring is not perfect, but it gets teams moving. After the first review, you can expand to a portfolio view. Some organizations add a compliance score or customer-commitment score for regulated workloads. Others include a churn risk multiplier for user-facing products.
Building your review team and roles
A good review balances detailed technical knowledge with cross-functional accountability. If you only invite platform engineers, you might get deep answers with little product context. If you only invite product stakeholders, you will miss the technical nuances that drive risk.
Define roles and expectations upfront: - Facilitator. Guides the meeting, manages time, resolves ambiguities, and ensures evidence is captured. Typically a cloud architect or platform lead. - Product owner. Provides business context, SLOs, user journeys, and ownership for priorities that trade off features versus risks. - Lead engineer. Walks through architecture details, deployment processes, and constraints. Supplies diagrams and code references. - Operations representative. Brings on-call experience, incident history, and runbook maturity. Demonstrates monitoring and alerting. - Security representative. Validates IAM, data protection, network segmentation, and secrets management choices. Clarifies policy expectations.
If you do not have a dedicated security engineer, assign someone accountable for interpreting the Security pillar questions and coordinating with governance functions. If you work in a mixed cloud environment, it can help to invite a peer from another platform to cross-pollinate practices. For example, if you also use Azure, compare how you handle identity and network controls across both clouds using the perspective from comparing Azure and AWS.
Give people time to prepare. Send the scope, a minimal architecture diagram, and a pre-read one or two days in advance. Clarify the expectation that answers will be supported by configuration evidence or IaC, not only by intent. Use calendar invites that split time by pillar, with a shared doc that already lists the questions you plan to cover.
Operational Excellence pillar: design for run, learn, and improve
Operational Excellence focuses on how you support your workload in production, from change management to observability and process learning. It is about how you operate, not just what you built. In practice, this pillar reveals whether the team can make changes safely, detect problems early, and learn quickly from incidents. It is also where you align engineering rituals with business outcomes through SLOs, runbooks, and feedback loops.
Start with change and deployment. Do you deploy using pipelines that enforce peer review, automated tests, and rollback mechanisms? If not, production stability depends on heroics. If you use infrastructure as code, is it the single source of truth, or do you rely on manual console changes for emergencies that never get reconciled? Establish a goal that all infrastructure is defined and reviewed as code, even if some legacy resources take time to migrate.
Next, assess observability. You should be able to answer three questions rapidly during an incident: what changed, where is the bottleneck, and who needs to act? This means centralized logging with queryable fields, metrics with clear thresholds and alarms, and traces that tie a symptom to a component. Use CloudWatch or a third-party solution consistently. Alert only on actionable signals. If an alarm does not lead to a decision, it is noise.
Finally, look at learning and improvement. Do you run incident reviews that focus on process and systems, not on blame? Do you write runbooks that are actually used and updated after incidents? Do you game day failure modes and rehearse disaster scenarios? The Operational Excellence pillar is about building muscles you use under stress.
Operational Excellence checklist: - All changes flow through versioned, peer-reviewed pipelines with automated tests and a rollback plan. - Infrastructure and configuration are managed as code. Manual changes are rare, documented, and reconciled. - You define service-level objectives and measure error budgets for key user journeys. - Centralized logs, metrics, and traces are implemented for all critical components, with alarms only on actionable thresholds. - Runbooks exist for top failure modes, with step-by-step actions, owners, and expected outcomes. - Incidents are reviewed with a focus on systemic improvement, and actions are tracked to closure. - Capacity planning is routine. You test scaling and throttling behaviors in non-production before events.
Example: a simple rollback plan with CodeDeploy blue-green for an ECS service could be documented as:
1. Create new task definition revision with image tag app:v2.
2. Deploy via CodeDeploy with traffic shifting set to 10 minutes linear.
3. Monitor 5xx rate and latency alarms for 15 minutes post-cutover.
4. If 5xx > threshold for 5 minutes, execute CodeDeploy rollback from console or CLI:
aws deploy stop-deployment --deployment-id d-ABC123 --auto-rollback-enabled
5. Investigate error budget burn and root cause in APM traces.
You can also codify operational checks. For example, use AWS Config rules to enforce that CloudWatch Alarms exist for key services. Simple enforcement reduces drift.
Security pillar: identity, data, network, and detection
Security in Well-Architected starts with identity and least privilege, extends through data protection and network segmentation, and finishes with detection and response. You are judged by how you prevent misuse and how you respond when, not if, something goes wrong. Reviews will quickly surface weak IAM policies, plaintext secrets, wide-open networks, and missing logs.
Begin with identity. Do you use IAM roles for workloads instead of long-lived access keys? Are policies scoped to the minimal actions and resources needed? Have you centralized human access with SSO and enforced MFA for all privileged users? If you operate multiple accounts, have you separated production and non-production, and applied service control policies appropriately?
Move to data protection. Are all data stores encrypted at rest with KMS keys that you control, not only AWS managed defaults? Do you have key rotation and clear key policies that limit administrative access? Are backups also encrypted and tested for restore? For data in transit, enforce TLS for all endpoints, and use private connectivity inside your VPCs.
Then network controls. Minimize the blast radius using security groups with least privilege rules, NACLs for coarse blocks, and VPC endpoints for private access to AWS services. Limit public subnets and avoid direct internet exposure where private access exists. Use WAF and Shield for internet-facing endpoints as appropriate. Ensure you log VPC Flow Logs for forensic needs.
Finish with detection and response. Centralize CloudTrail across accounts, aggregate in a dedicated logging account, and protect those logs from tampering. Enable GuardDuty for threat detection. Implement AWS Config rules for continuous assessment. Practice incident response with playbooks that cover credential compromise, data exfiltration, and ransomware scenarios.
Security checklist: - Human access uses SSO with MFA. Machine access uses IAM roles, not long-lived keys. - IAM policies are least privilege, resource-scoped, and regularly reviewed. Access Analyzer is used to find unintended access. - All data at rest is encrypted with KMS CMKs where needed, with rotation and clear key policies. Secrets are stored in Secrets Manager or Parameter Store with encryption. - All data in transit uses TLS. For internal services, enforce minimum TLS versions and private access. - Network design isolates tiers and environments. Security groups deny by default, and VPC endpoints limit internet exposure. - CloudTrail is enabled organization-wide and delivered to a protected S3 bucket. GuardDuty and Config are enabled, with findings triaged. - Incident response runbooks exist, are tested, and include contacts, containment steps, and notification criteria.
Example IAM policy anti-pattern and fix: - Anti-pattern: An application role with "Action": "", "Resource": "", justified by "it needs to access a lot". - Fix: Start from AWS managed policies for the relevant service, then trim to actions used. Scope resources with ARNs. If dynamic resources are created, use condition keys and naming conventions rather than wildcards.
If you want to go deeper on identity design, visit the focused guide to AWS security, identity, and access management practices. Use that alongside your review to convert findings into concrete policy and account structure changes.
Reliability pillar: recover, scale, and manage quotas
Reliability in Well-Architected is about the ability of a workload to perform its intended function correctly and consistently when expected. That includes how you plan for failure, how you recover, and how you manage capacity and quotas. In AWS, this maps to multi-AZ designs, resilient state management, automated recovery, and tested backup and restore.
Start with failure isolation. For stateful services like RDS or ElastiCache, run in multiple Availability Zones where supported. For compute, spread instances across subnets in at least two AZs behind load balancers. For event processing, use services with dead-letter queues and retry semantics. Avoid single points of failure like a single NAT Gateway in a region-wide architecture if you have high east-west traffic.
Backups and restore matter more than backups alone. Define RPO and RTO for each data store: - RPO, Recovery Point Objective, is how much data loss you can tolerate. For example, RPO 5 minutes means your restore must recover to a point not older than 5 minutes before failure. - RTO, Recovery Time Objective, is how long you can be down during recovery. For example, RTO 30 minutes means full service recovery must complete within 30 minutes.
Ensure automated backups with tested restore drills. For RDS, test point-in-time recovery to a new instance in a separate subnet group. For EBS, validate that snapshots can be used to restore volumes and reattach instances. For S3, use versioning and lifecycle rules, and consider replication for critical buckets.
Quota and capacity management is often missed. Understand service quotas for your regions and request increases before events. Implement autoscaling for compute and concurrency limits for serverless. Design for graceful degradation when downstream services throttle. Build idempotent consumers so retries do not create cascades.
Reliability checklist: - Compute and stateful services are deployed across multiple AZs with health checks and automated failover. - Backups are automated, encrypted, and regularly tested for restore. RPO and RTO are defined per workload component. - Event-driven components use retries with backoff and dead-letter queues, with monitoring on DLQ depth. - Quotas and limits are tracked, and increases are requested proactively. Workload handles throttling gracefully with exponential backoff and jitter. - Infrastructure has self-healing mechanisms, such as autoscaling and replacement of unhealthy instances. - Chaos and game day exercises validate failure assumptions, including dependency outages and credential compromise.
A simple example of exponential backoff with jitter in pseudocode:
base = 100ms
for attempt in 1..5:
sleep = random(0, base * 2^attempt)
try request
if success: break
Using jitter avoids synchronized retries that can worsen an outage.
Performance Efficiency pillar: right place, right size, right time
Performance Efficiency centers on how you use computing resources efficiently to meet requirements and maintain efficiency as demand changes. It is not about chasing the biggest instance type, but about selecting the right services, managing load patterns, and measuring bottlenecks with evidence. AWS gives you managed services that can absorb burst and scale linearly, but only if your design and limits align with reality.
Start with service and architecture choices. If your workload is event-driven or has variable traffic, consider managed services like Lambda, DynamoDB with on-demand capacity, or SQS decoupling. If you run microservices, use load balancing and connection pooling to avoid hot spots. If you rely on EC2, pick instance families aligned to your needs, such as compute-optimized for CPU-bound workloads or memory-optimized for in-memory caches.
Next, measure and tune. Use profiling and APM tools to find application-level bottlenecks before scaling infrastructure. Add caching where data access patterns are read-heavy. Use content delivery networks to move content closer to users. Size database connections and thread pools based on observed limits, not defaults. Monitor tail latencies, not just averages.
Finally, automate scaling and adapt. Set autoscaling policies with sensible cooldowns and predictive signals when possible. For Lambda, manage concurrency reservations for critical functions to avoid cold start contention during spikes. For DynamoDB, use adaptive capacity and watch throttling metrics. For queues, aim for steady consumers that avoid bulk spikes that cause downstream overload.
Performance Efficiency checklist: - Architecture choices align with workload patterns. Event-driven and bursty workloads use managed services where appropriate. - Bottlenecks are measured at application and infrastructure layers, including p95 and p99 latencies. - Caching strategies are documented and implemented, including CDN for static and dynamic content where viable. - Autoscaling is enabled and tuned with appropriate min, max, and cooldowns. Serverless concurrency limits are defined and tested. - Thread pools, DB connections, and timeouts are sized based on evidence. Retry strategies use backoff with jitter. - Load testing is part of release preparation, with tests that reflect realistic traffic patterns.
If you are evaluating managed compute options, it can help to study trade-offs and operating models in serverless architectures on AWS. Serverless can improve efficiency by shifting undifferentiated heavy lifting to AWS, provided you design for concurrency and cold start realities.
Cost Optimization pillar: spend where it matters, eliminate waste
Cost Optimization is about controlling spend without compromising reliability or performance. The big wins come from aligning capacity with demand, rightsizing, using the right pricing model, and eliminating orphaned resources. The smaller but important wins come from governance, cost visibility, and engineering practices that treat cost as a dimension of quality.
Begin with visibility. Enable AWS Cost Explorer, define cost and usage reports, and enforce tagging for cost allocation. Set budgets and alerts for key dimensions. Give product owners dashboards that map spend to value, such as cost per transaction. If you lack good tags, start with a tag policy and a plan to improve coverage gradually.
Rightsizing and scheduling come next. Review EC2, RDS, and EKS nodes for underutilization. Use instance families and sizes that match observed CPU and memory. For development and staging, schedule instances to stop out of hours. For Lambda and Fargate, watch duration and memory allocations to avoid over-provisioning. Eliminate unattached EBS volumes, idle load balancers, and old snapshots.
Pick the right pricing model. Use Savings Plans or Reserved Instances for steady-state compute. Buy based on your conservative baseline, then use on-demand and spot for burst and batch. Evaluate Graviton adoption where supported for better price-performance. For storage, apply lifecycle policies to move infrequently accessed data to cheaper tiers.
Cost Optimization checklist: - Cost allocation tags are defined, enforced, and used in Cost Explorer and reports. Owners can see cost per product or environment. - Budgets and alerts are configured for key services and teams. Cost anomalies are reviewed and acted upon. - Instances, databases, and containers are right-sized based on observed utilization. Development resources are scheduled to stop when idle. - Savings Plans or RIs cover baseline compute. Spot is used for fault-tolerant workloads. Graviton adoption is evaluated for eligible services. - Storage lifecycle policies and data retention rules prevent unbounded growth. Snapshots and buckets are cleaned up regularly. - Engineering practices consider cost in design, including query optimization and efficient data serialization.
For a deeper walkthrough of specific tactics and how to operationalize them with teams and tooling, use the step-by-step guidance in AWS cost optimization strategies. If your organization is new to AWS, start with foundational skills in AWS fundamentals before attempting portfolio-wide cost governance.
Sustainability pillar: efficiency, hardware choice, and data lifecycle
Sustainability in the Well-Architected Framework focuses on minimizing the environmental impacts of your workload by improving energy efficiency and making informed choices that reduce compute and data resource needs. While AWS operates data centers efficiently, your architectural and operational choices determine the resources your workload consumes. This pillar helps you reduce waste and align engineering with broader sustainability goals.
Start with efficient resource use. Right-size compute and memory, and use serverless or managed services that share resources efficiently. Keep hot paths lean, reduce idle time, and collapse layers that add latency and energy without business value. Avoid over-replication of data and unnecessary conversions that multiply storage and compute.
Then consider hardware and region choices. Where possible, adopt energy-efficient instance types such as those based on AWS Graviton. Evaluate the carbon intensity of regions alongside latency and regulatory requirements. If you run data heavy workloads, optimize data formats and compression to minimize transfer and storage energy.
Data lifecycle management is a powerful lever. Set clear retention policies, move cold data to archival tiers, and delete data that no longer serves a purpose. Prefer query patterns that avoid full scans over time-series data. For analytics, use partitioning and predicate pushdown to reduce scan size and compute.
Sustainability checklist: - Compute is right-sized and scales with demand. Idle resources are eliminated using schedules and autoscaling. - Serverless and managed services are adopted where they reduce resource consumption for the workload. - Energy-efficient hardware options, such as Graviton-based instances, are evaluated and adopted where supported. - Region selection considers carbon intensity alongside legal and latency requirements. - Data lifecycle policies are defined and enforced. Data formats and compression reduce storage and transfer. - Workload metrics include resource efficiency indicators, such as cost and energy proxy per transaction.
Sustainability is integral to cost and performance too. You will often find that the same practices that reduce spend and improve performance also reduce environmental impact. Use the Sustainability pillar to make that explicit in your design decisions and dashboards.
Using the AWS Well-Architected Tool effectively
The AWS Well-Architected Tool is your system of record for the review. It standardizes questions, highlights risks, and tracks improvements. Used well, it saves time and helps with audits. Used poorly, it becomes a checkbox that diverges from the truth.
Here is a practical way to work with the tool: - Create the workload with a clear, stable name. Add a description that includes business context and the in-scope accounts and regions. - Share the workload with collaborators across accounts if needed. Limit who can change answers to the facilitator and lead engineer. - Move pillar by pillar. For each question, discuss aloud, then record the answer and rationale in the notes. Add links to evidence, such as repository paths, runbook pages, or dashboard URLs. - When the tool flags a high risk, rephrase it as a concrete improvement in your own tracker, with acceptance criteria. The tool is not a task manager, so do not hide work inside it.
The tool also supports lenses for specific domains, such as serverless or data analytics. Use lenses when your workload heavily depends on those patterns. However, do not let lenses displace the core six pillars, which must be consistently applied across workloads. For a reference on functionality and sharing, see the AWS user guide at https://docs.aws.amazon.com/wellarchitected/latest/userguide/what-is-wellarchitected-tool.html.
Link the tool to your working artifacts. For example, include the Well-Architected workload ID in your improvement plan spreadsheet and in tickets. This lets you correlate improvements back to a review instance, which is useful for compliance audits and trend analysis across multiple reviews.
Common findings and anti-patterns by pillar
The same issues appear repeatedly across reviews. Recognizing them early lets you ask sharper questions and propose tested fixes. Use the list below as both a pre-mortem and a post-review checklist to ensure you captured the expected risks with enough detail.
Operational Excellence: - No rollback plan or rollback is manual and untested. Fix by implementing deployment strategies with automated rollback and practice them. - Logs are siloed or missing correlation IDs. Fix by standardizing log formats and pushing to a centralized log platform with trace correlation. - Incident reviews are ad hoc and do not produce actions. Fix by adopting a light, consistent template and assigning follow-through owners.
Security: - Developers use long-lived access keys. Fix by moving to IAM roles and short-lived credentials via SSO. - KMS key policies grant excessive admin rights to the same team that uses the keys. Fix by separating key management roles and limiting grants. - Public S3 buckets or overly permissive bucket policies exist. Fix by blocking public access at the account level and reviewing exceptions.
Reliability: - Single NAT Gateway in a high-traffic architecture. Fix by adding NAT Gateways per AZ for resilience and throughput, or redesign with VPC endpoints. - No tested database restore. Fix by scheduling restore drills with success criteria and documenting steps in runbooks. - No alarm on DLQ depth. Fix by adding alarms, dashboards, and a runbook to triage and reprocess messages.
Performance Efficiency: - Over-sized instances running at 5 percent CPU. Fix by rightsizing, using autoscaling, and profiling app performance. - Unbounded downstream calls causing timeouts under load. Fix by adding timeouts, circuit breakers, and bulkhead isolation. - Cold start sensitivity in Lambda due to large packages. Fix by trimming dependencies, using provisioned concurrency for critical paths.
Cost Optimization: - Missing tags prevent spend accountability. Fix by defining mandatory tags and enforcing them with tag policies and automation. - Idle load balancers and unattached EBS volumes. Fix by periodic cleanup jobs and alerts for idle resources. - No Savings Plans for steady compute. Fix by buying a conservative baseline and monitoring coverage.
Sustainability: - Excessive data retention with no legal or business need. Fix by defining retention policies and automating deletion. - Inefficient data formats causing full scans. Fix by adopting columnar formats with partitioning and compression. - Instances not using energy-efficient options without reason. Fix by evaluating Graviton migration for supported workloads.
A compact mapping for quick reference:
| Pillar | Common issue | Impact | Quick win | Longer fix |
|---|---|---|---|---|
| Security | Long-lived access keys | Credential leak risk | Rotate and disable keys, move to role assumption | Enforce SSO and federation, remove key use by policy |
| Reliability | Untested DB restore | Prolonged outage | Schedule restore drill | Automate restore tests in pipeline |
| Cost | Idle load balancers | Waste | Delete idle ALBs/NLBs | Automate idle resource detection |
| Performance | Oversized instances | Waste and inefficiency | Rightsize | Autoscaling and app profiling |
| Ops | No rollback plan | Slow recovery | Document and test basic rollback | Adopt blue-green or canary deployment |
| Sustainability | Excessive retention | Cost and footprint | Apply lifecycle rules | Redesign data model and retention governance |
Prioritizing remediations with a simple scoring model
After the review, you may have dozens of findings. Choosing what to do first is where many reviews stall. Use a simple, transparent scoring model to rank items, then revisit the ranking with business context. The goal is to drive agreement quickly and move to execution.
Define a Risk Priority Number, RPN, as: - Severity, S, 1 to 5. How bad is the impact if this risk materializes? - Likelihood, L, 1 to 5. How likely is it to occur in the next 6 to 12 months? - Effort, E, 1 to 5. How much engineering time does it take to fix? Use inverse weighting so lower effort increases priority.
Compute:
RPN = S * L * (6 - E)
This way, low effort items get a boost. For example: - Unencrypted S3 bucket for logs. S=4, L=3, E=2. RPN = 43(6-2) = 48. - Database restore drill. S=5, L=3, E=3. RPN = 53(6-3) = 45. - Switch EC2 to Graviton. S=2, L=3, E=4. RPN = 23(6-4) = 12.
Now sort by RPN and draw a line where your team capacity ends for the quarter. Share the ranked list with product leadership. If a lower RPN item satisfies a regulatory requirement or an upcoming audit, move it up explicitly and document the reason. The transparency reduces debate time and builds trust.
You can add a dependency field to capture sequencing. For example, you may need to establish tagging standards before you can apply automated cost policies. Mark those relationships so you do not schedule blocked work. Track decisions in the improvement plan, not in people’s heads.
Turning findings into a workload improvement plan
A Well-Architected Review is only as good as the improvements that follow. Convert each finding into a concrete improvement with an owner, timeline, and acceptance criteria. Keep the plan lean and visible to both engineering and product stakeholders. Use your existing ticketing system, but also maintain a summary view that is easy to discuss in steering meetings.
Use a simple template: - Finding. A concise statement of the risk, with source evidence. - Pillar. One of the six pillars. - Severity. 1 to 5. - Owner. One accountable person or team. - Due date. Target completion date. - Dependencies. Prior items or external decisions. - Acceptance criteria. Objective, verifiable conditions that demonstrate closure.
Example improvement plan table:
| ID | Finding | Pillar | Severity | Owner | Due date | Dependencies | Acceptance criteria |
|---|---|---|---|---|---|---|---|
| IMP-01 | CloudTrail not enabled in all accounts | Security | 5 | Platform | Aug 15 | Org SCP change approved | CloudTrail enabled org-wide, logs delivered to central S3 with MFA delete, verified in all accounts |
| IMP-02 | No tested RDS point-in-time restore | Reliability | 5 | DB Team | Aug 30 | None | Restore performed to new instance, read-only check passes, runbook updated |
| IMP-03 | 30 percent of EC2 instances under 10 percent CPU | Cost | 3 | App Team A | Sep 10 | Tag coverage > 90 percent | Rightsize 80 percent of candidates, 20 percent exception list approved |
| IMP-04 | No rollback procedure for ECS service | Ops | 4 | App Team B | Aug 22 | None | Blue-green deploy with rollback tested in staging and production once |
Review the plan weekly, not monthly. Celebrate closures and call out blockers. If you need to upskill the team to tackle an area, invest early. You can accelerate capability building with structured learning and hands-on mentorship. If you are building toward a career in this space, explore how the Refonte Cloud Engineer Program blends guided projects and internships to make these practices second nature.
Review cadence, triggers, and continuous improvement
The best teams do not treat a Well-Architected Review as a one-time event. They run lightweight check-ins on a cadence and trigger focused mini-reviews when certain events happen. Over time, this habit dramatically reduces the size and surprise of full reviews.
Set a baseline cadence. For critical customer-facing services, aim for a full review annually, with quarterly check-ins on unresolved items and any new risks. For internal services, semiannual may be sufficient. Tie check-ins to roadmap planning so improvements can be funded and staffed alongside features.
Define triggers for ad hoc reviews: - Major architecture changes, such as moving from EC2 to EKS or adding a new data store. - Scale events, like big marketing campaigns or new region launches. - Compliance or customer due diligence requests. - Incident postmortems that reveal systemic weaknesses.
Automate guardrails where possible. Use AWS Config, CloudFormation hooks, and organizational policies to prevent or flag high-risk patterns at provisioning time. Add pipeline checks for critical controls, like ensuring that all S3 buckets are private by default or that security groups with 0.0.0.0/0 ingress require an exception process.
Integrate measurement with your observability platform. Track trends in error budgets, incident counts, and high-risk findings over time. Publish a small, stable set of metrics that executives can understand, such as percentage of workloads with tested restores or tag coverage above 90 percent. This keeps the conversation on progress, not on isolated anecdotes.
Running the workshop: facilitation tactics that work
The quality of your Well-Architected Review depends as much on facilitation as on technical detail. A well-run session produces alignment and actionable outcomes. A rushed or confrontational session produces defensive answers and little follow-through. Approach the workshop as a collaborative design review with production reality in mind.
Time-box each pillar, but leave room for deeper dives when evidence is missing or a high-risk pattern emerges. When an answer is unclear, ask for a demonstration in the console or via IaC. For example, to verify CloudTrail configuration, ask the operator to show the organization settings and the destination bucket policy. Evidence reduces back-and-forth later and avoids misunderstandings.
Use a parking lot for deep topics you cannot resolve on the spot. If IAM policy design becomes a rabbit hole, note it as a follow-up with a smaller group. Keep the main session on track to cover all pillars. For each risk you capture, immediately ask for an owner and a next step. Do not leave ownership ambiguous.
Close with a recap that includes the top five findings by severity, the owners, and the expected timeline for the improvement plan draft. Confirm the date for the follow-up meeting where you will finalize prioritization and accept the plan. This ritual builds momentum and ensures the session leads to action.
Evidence collection and examples to speed up your review
Evidence reduces argument and saves time. Prepare lightweight scripts and queries that quickly answer common questions. A small investment here can cut hours from the workshop and avoids subjective debate.
Useful examples: - List S3 buckets that are public or allow public ACLs:
aws s3api list-buckets --query "Buckets[].Name" --output text | tr '\t' '\n' | while read b; do
policy=$(aws s3api get-bucket-policy-status --bucket "$b" --query PolicyStatus.IsPublic --output text 2>/dev/null)
acl=$(aws s3api get-bucket-acl --bucket "$b" --query "Grants[?Grantee.URI=='http://acs.amazonaws.com/groups/global/AllUsers']" --output text 2>/dev/null)
if [ "$policy" = "True" ] || [ -n "$acl" ]; then echo "Public: $b"; fi
done
- Verify CloudTrail org-level logging:
aws organizations describe-organization >/dev/null 2>&1 && \
aws cloudtrail list-trails --query "Trails[?IsOrganizationTrail==\`true\`]"
- Find unattached EBS volumes older than 30 days:
aws ec2 describe-volumes --filters Name=status,Values=available \
--query "Volumes[?CreateTime < \`$(date -d '30 days ago' -Ins --utc | sed 's/+00:00/Z/')\`].[VolumeId,Size,CreateTime]" --output table
Dashboards to prepare: - Cost Explorer views by tag keys such as Application, Environment, and Owner. Show monthly spend, daily anomalies, and top services. - CloudWatch dashboards for p95 latency, error rates, and key dependency health. Add DLQ depth and retry rates where event-driven. - A diagram that clearly shows AZ distribution, public and private subnets, and VPC endpoints.
The goal is not to automate the entire review, but to anchor answers in facts you can verify. Bring these scripts to the session and run them as needed. You will quickly separate assumptions from reality.
Checklists you can reuse, by pillar
Centralize your checklists and reuse them across workloads. Tailor wording to your environment so teams can apply them without translation.
Operational Excellence quick checklist: - Pipelines enforce peer review, automated tests, and rollback. Deploys are repeatable. - Infra as code is the source of truth. No snowflake resources. - SLOs and error budgets exist for main user journeys. Alarms align to SLOs. - Logs, metrics, and traces are centralized with actionable alerts. Runbooks are current. - Incident reviews produce actions tracked to closure. Gamedays run at least twice per year.
Security quick checklist: - SSO with MFA for humans, roles for machines. No long-lived access keys. - Least privilege IAM with resource scoping. Access Analyzer findings triaged. - KMS used for encryption at rest, with key policies and rotation. Secrets in managed stores. - TLS enforced in transit. Private access for internal services. - VPC segmentation and VPC endpoints. GuardDuty, CloudTrail, and Config enabled and monitored. - Tested incident response for credential leaks and data exfiltration.
Reliability quick checklist: - Multi-AZ deployments and health checks. Self-healing in place. - Tested backups and restores. RPO and RTO defined and met. - Retry and DLQ patterns used. Alarms on DLQ depth. - Quotas monitored and increased proactively. Graceful degradation under throttle. - Chaos tests and gamedays validate assumptions.
Performance Efficiency quick checklist: - Architecture matches pattern. Use managed services where fit. - Profiling and APM in place. Tail latencies tracked. - Caching and CDN strategies implemented. Hot paths optimized. - Autoscaling and concurrency limits tuned. Timeouts and retries right-sized. - Load tests reflect real traffic and data.
Cost Optimization quick checklist: - Cost allocation tags enforced. Budgets and anomaly alerts active. - Rightsizing and schedules applied. Idle resources cleaned. - Savings Plans or RIs cover baseline. Spot for fault-tolerant tasks. - Storage lifecycle and retention policies enforced. - Engineers consider cost in design and reviews.
Sustainability quick checklist: - Right-size and autoscale. Idle eliminated. - Efficient hardware options considered. Serverless where fit. - Regions selected with sustainability in mind. - Data lifecycle defined. Efficient formats and compression used. - Efficiency metrics included alongside performance and cost.
If you want to build foundational habits that make these checklists routine, spend time with the fundamentals in AWS architecture basics and core services. Improved fundamentals shorten reviews and accelerate fixes.
Explore the silo
- Visit the cloud skills and architecture hub for foundational guides and all cloud topics.
- Deepen your base with AWS fundamentals and core services.
- Strengthen your guardrails with AWS security and IAM essentials.
- Cut waste with practical AWS cost optimization.
- Evaluate patterns in serverless architectures on AWS.
- Compare platforms in Azure vs AWS for teams and workloads.
- Plan your learning path via cloud certifications and career paths.
- Track what's changing in DevOps trends, tools, and career guide.
- Level up monitoring with learn Prometheus for free.
FAQ
Q: What is an aws well architected review and how is it different from an audit? A: It is a collaborative assessment of a single workload against AWS best practices across six pillars. The goal is to identify risks and improvements, not to assign blame or pass-fail grades. Unlike audits, a review produces an improvement plan owned by the product and engineering team and is intended to be repeated as your workload evolves.
Q: How long does a Well-Architected Review take? A: Plan 2 to 6 hours of workshops, typically split across two sessions, plus time to gather evidence and document improvement items. The duration depends on workload complexity and team readiness. Subsequent reviews take less time once you build a baseline and automate guardrails.
Q: Do we need to use the AWS Well-Architected Tool? A: You do not have to, but it standardizes questions, flags risks consistently, and serves as a record of the review. It is useful for repeatability and audits. Many teams use the tool for answers and risks, then manage remediation in their usual ticketing systems with cross-links.
Q: Which workload should we review first? A: Choose a high-impact workload with active development and clear ownership. Use a simple scoring model based on business impact, user exposure, and change velocity. Your first review should generate momentum and improvements, so avoid a legacy system that no one can change quickly.
Q: What evidence should we prepare before the workshop? A: Prepare an architecture diagram, CloudTrail and Config summaries, IAM role and policy lists, backup and restore configurations, CloudWatch dashboards and alarms, and Cost Explorer views with tag coverage. Having these at hand shortens discussion and anchors answers in facts.
Q: How do we prioritize the findings? A: Use a simple scoring model that multiplies severity and likelihood and boosts low-effort items. Then adjust for regulatory or customer commitments and for dependencies. Publish the prioritized list with owners and due dates so stakeholders can challenge and confirm the plan.
Q: How often should we repeat the review? A: Aim for a full review annually for critical services, with quarterly check-ins on open items and new risks. Trigger focused mini-reviews when you make major architecture changes, prepare for scale events, or complete significant incident investigations.
Q: Where can my team build the skills to run and act on reviews effectively? A: Make the framework part of onboarding, share checklists, and practice on non-critical workloads first. For more structured development with hands-on projects and mentorship, see the Refonte Cloud Engineer Program, which focuses on operational excellence, security, cost, and reliability as everyday engineering practices.
