AWS Security and IAM: Least Privilege, SCPs, and Blast Radius Design
AWS security is not a product you buy, it is a set of design decisions you make before a single workload runs. The accounts you create, the identities you issue, the guardrails you enforce, and the ways you contain failure all shape whether one compromised credential becomes a footnote or a headline. This pillar page walks through the identity, org, encryption, and detection primitives that carry the weight of most real breaches and audits, and shows you how to compose them into a defensible architecture. You will see the patterns that pass audits, the mistakes that show up in postmortems, and the concrete steps to move from a single sprawling account to a bounded, observable multi-account estate. Read it end to end if you are designing a landing zone, or jump to sections when you need to make a specific decision.
The mental model: identity, boundaries, and blast radius
Most AWS security incidents are not exotic. They are variations on three themes: an identity did something it should not have been allowed to do, a boundary was missing or misconfigured, and the blast radius was larger than anyone realized. If you internalize those three ideas, you can evaluate almost any AWS security decision on the fly.
Identity is who or what is making the API call. In AWS that is an IAM user, an IAM role assumed by a human via federation, a role assumed by a workload (EC2, Lambda, ECS task, EKS pod), or a root user you should almost never use. Every request to AWS carries an identity, and every request is evaluated against a chain of policies. If you cannot answer "which identity did this and what policies applied" in under a minute, your identity design is already too complex.
Boundaries are the walls that limit what an identity can do even if its own policy says otherwise. Permission boundaries cap what an IAM principal can be granted. Service control policies (SCPs) cap what any principal in an account or organizational unit can do. Resource policies decide who can touch a specific resource. Session policies narrow a role's permissions at assume-time. These boundaries compose, and the effective permission is always the intersection, minus any explicit deny.
Blast radius is what an attacker or a mistake can reach once inside. A single account with every workload in it has a huge blast radius. A per-workload account with tight SCPs and no cross-account trust has a small one. The whole point of multi-account design, VPC segmentation, KMS key scoping, and least-privilege IAM is to make the blast radius of any single failure small enough that you can absorb it, detect it, and recover without a company-wide incident.
Everything else in this article is machinery for those three ideas. If you are new to the platform, spend an afternoon with the AWS fundamentals overview before diving deeper, because IAM assumes fluency with regions, services, and the shared responsibility model.
IAM users, roles, and the case against long-lived credentials
IAM users are the oldest identity primitive in AWS and, for human access, they are almost always the wrong choice today. A user has long-lived credentials: a password, and optionally access keys that are just an ID and a secret with no expiry. If either leaks, the attacker has persistent access until someone notices and rotates. The number of public breaches traced to a leaked access key in a Git repo or a laptop backup is embarrassing.
The modern pattern is federation plus roles. You keep your identities in an identity provider you already trust, whether that is AWS IAM Identity Center (formerly SSO), Okta, Entra ID, or Google Workspace. Humans sign in there, get a short-lived session, and assume an IAM role in the target AWS account. The role has the permissions, not the human. Credentials expire in an hour or a few hours. When someone leaves the company, you disable them in one place and their AWS access dies with the next session refresh.
For workloads the story is similar. EC2 instances get an instance profile that maps to a role. Lambda functions have an execution role. ECS tasks have a task role. EKS pods use IAM roles for service accounts (IRSA) or the newer EKS Pod Identity. Every one of these mechanisms delivers temporary credentials via the instance metadata service or the SDK credential provider chain. You never bake an access key into an AMI, a container image, or a config file. If you find yourself typing aws configure with a static key on a server, stop and redesign.
The remaining legitimate uses of IAM users are narrow: a break-glass account you keep offline for disaster recovery when your identity provider is down, a small number of programmatic integrations with vendors that genuinely cannot do OIDC or role assumption, and some CI systems on older setups. Even those integrations should move to OIDC federation where possible. GitHub Actions, GitLab, CircleCI, and Buildkite all support OIDC into AWS today, which means your pipelines assume a role with a short-lived token instead of storing an access key in a secret.
If you must have an IAM user, enforce MFA, rotate keys on a schedule you actually follow, and set a permission boundary so the blast radius is limited even if the credential leaks. And log every use to CloudTrail so you can prove when it was last touched.
Policies: identity, resource, session, and the evaluation order
An IAM policy is a JSON document that grants or denies actions on resources under conditions. There are several kinds and they interact in a specific order that trips people up constantly.
Identity policies are attached to a user, group, or role and describe what that principal can do. Resource policies are attached to a resource (an S3 bucket, a KMS key, a Lambda function, an SNS topic) and describe who can touch it. Permission boundaries are attached to a principal and cap the maximum permissions that principal can ever have, regardless of what identity policies say. Session policies are passed at assume-role time and further narrow a session. SCPs sit at the organization level and cap what any principal in the account can do.
The evaluation logic, roughly: an explicit deny anywhere always wins. Otherwise, the request must be allowed by at least one applicable policy in each category that applies. For a call within the same account, that is usually the identity policy and any SCP. For cross-account calls, both accounts must allow it: the source account's identity policy must allow the action, and the target account's resource policy must allow the principal. When a permission boundary is attached, the effective identity permissions are the intersection of the identity policy and the boundary.
A concrete example. Suppose a developer role has an identity policy allowing s3:* on arn:aws:s3:::reports-*. A permission boundary on that role allows only s3:GetObject and s3:ListBucket. The effective permission is s3:GetObject and s3:ListBucket on arn:aws:s3:::reports-*, because the boundary caps the identity policy. Now add an SCP on the account that denies s3:* unless the request comes from a specific VPC endpoint. The developer's calls from a laptop over the public internet fail, even though the identity policy and boundary both allowed them, because the SCP's deny is absolute.
Writing good policies means being specific about actions, resources, and conditions. Resist the urge to write "Action": "*" or "Resource": "*" outside of narrowly scoped admin roles. Use condition keys like aws:PrincipalTag, aws:ResourceTag, aws:SourceIp, aws:VpcSourceIp, aws:MultiFactorAuthPresent, and aws:RequestedRegion to add context. A policy that grants ec2:TerminateInstances only when aws:ResourceTag/Environment equals dev is dramatically safer than one that grants it globally.
Permission boundaries: the delegation pattern that actually scales
Permission boundaries solve a specific organizational problem: you want to let developers create IAM roles for their workloads without giving them the power to grant themselves admin. Without boundaries, any developer with iam:CreateRole and iam:AttachRolePolicy can trivially escalate to admin by creating a role with AdministratorAccess and assuming it.
The pattern is: create a boundary policy that describes the maximum permissions any workload role in your account should have, for example, read/write to specific S3 prefixes, access to specific DynamoDB tables, publish to specific SNS topics, and nothing more. Then grant developers permission to create roles only if the new role has that boundary attached. The IAM policy that grants role creation uses a condition like "iam:PermissionsBoundary": "arn:aws:iam::123456789012:policy/DeveloperBoundary".
Now developers can iterate on IAM roles for their services without opening tickets, and the worst case is bounded. Even if they attach AdministratorAccess to their new role by mistake, the effective permissions are still capped by the boundary. This is one of the few AWS features that genuinely changes how fast a security team can safely delegate.
Boundaries also make audits easier. Instead of reviewing every role's identity policy, you review the boundary once. Any role with that boundary attached is provably incapable of exceeding it. Combine boundaries with a naming convention (role/workload/*) and tag-based access control, and you have a delegation model that survives contact with a growing engineering org.
The downside is that boundaries are one more thing to reason about, and forgetting to attach one is a silent failure. Enforce their presence with SCPs or with a Config rule that alerts on any role in the workload path without the required boundary.
AWS Organizations, OUs, and service control policies
A single AWS account is fine for a hobby project. For anything with real users, real data, or real regulatory scope, you want multiple accounts under an organization. The reasons pile up quickly: blast radius isolation, blast radius isolation, easier billing separation, per-environment quota headroom, cleaner IAM scoping, and simpler audit narratives.
A typical landing zone has a management account (with no workloads, only Organizations and billing), a log archive account, an audit account, a shared services account for things like DNS and CI, and then per-environment or per-team workload accounts. Organizational units (OUs) group accounts for policy application: Prod, Non-Prod, Sandbox, Security, Suspended. You attach SCPs to OUs, not usually to individual accounts, so that policy is inherited and consistent.
SCPs are guardrails, not grants. They can only deny or restrict, they never expand permissions. A common set of SCPs to start with:
- Deny disabling of CloudTrail, GuardDuty, Config, and Security Hub. These are your evidence and detection layers. Nobody in a workload account should be able to turn them off.
- Deny use of the root user for anything other than the specific actions that require it.
- Deny actions in regions you do not use. If your business runs in
us-east-1andeu-west-1, deny all actions in every other region. This is one of the highest-value SCPs you can write, because a lot of crypto-mining attacks spin up instances in obscure regions to avoid detection. - Deny creation of IAM users in workload accounts. Force everyone to federated roles.
- Deny modification of resources tagged
aws:cloudformation:ormanaged-by:platformoutside the deployment pipeline. - Require encryption on S3 buckets and EBS volumes at creation time.
SCPs apply to everything in the account except the management account itself, which is a reason to keep the management account nearly empty. If an SCP denies an action, no identity policy, boundary, or resource policy can override it. That is what makes SCPs the strongest guardrail AWS offers.
Where SCPs fit in a broader architecture is a topic the AWS Well-Architected security pillar guide treats in more depth, alongside the operational, reliability, and cost pillars. Read those together because security decisions constantly trade against other pillars.
Multi-account patterns and the shape of a landing zone
The specific account topology depends on your size, but a few patterns hold up well.
For a small company (under fifty engineers), a reasonable starting shape is: management, log-archive, audit, shared-services, three workload accounts per environment (dev, staging, prod), and a sandbox per engineer or team. That is roughly ten to fifteen accounts. AWS Control Tower or a hand-rolled Terraform landing zone can create and govern them.
For a larger org, you go per-team or per-product per-environment: payments-prod, payments-staging, payments-dev, catalog-prod, and so on. This can easily grow to hundreds of accounts. At that scale you must automate account creation, baseline configuration, and IAM Identity Center permission set assignment. Any manual step in the account vending process becomes a bottleneck and a source of drift.
Networking across accounts is usually done with Transit Gateway or VPC peering, with a hub account owning the shared network. Shared services like Route 53 private hosted zones, ACM Private CA, and internal package registries live in the shared-services account and are consumed by workload accounts via cross-account resource sharing (RAM) or resource policies.
Log aggregation flows to the log-archive account. CloudTrail organization trails, VPC flow logs, Config snapshots, and application logs all land in an S3 bucket in that account, with a bucket policy that only allows write from organization principals and read only from the audit account. That bucket is the single source of truth for security investigations, and it should have object lock and MFA delete configured so nobody, including you, can erase evidence.
The audit account holds Security Hub as delegated administrator, GuardDuty as delegated administrator, and any third-party CSPM tools. Read-only cross-account roles let security engineers pivot into workload accounts for investigation without needing standing access.
KMS: key hierarchy, grants, and the encryption defaults you should set
Encryption at rest on AWS is usually a checkbox, but the interesting decisions are about keys, not ciphers. KMS is the service that manages keys, and every service that supports encryption at rest (S3, EBS, RDS, DynamoDB, Secrets Manager, SNS, SQS, Lambda environment variables, and dozens more) integrates with it.
There are three flavors of KMS key: AWS-owned keys (managed by AWS, invisible to you, free), AWS-managed keys (managed by AWS but visible in your account, still free, one per service like aws/s3), and customer-managed keys (CMKs, which you create and pay for at $1/month each plus API calls). The rule of thumb: if you need to control who can decrypt, if you need a key policy of your own, if you need to rotate on your own schedule, or if you need to cross-account share, use a CMK. Otherwise the managed keys are fine.
Key policies are the resource policy on a KMS key and they are the primary authorization mechanism. Unlike most other AWS resources, IAM policies alone cannot grant access to a KMS key; the key policy must also allow the principal. This is intentional and prevents an over-broad IAM policy from silently unlocking every key in the account.
A common key policy pattern gives the key administrator role (usually a security team role) permission to manage the key, and gives specific workload roles permission to encrypt and decrypt. Grants are a lightweight alternative for temporary, programmatic access, often used by AWS services on your behalf.
Encryption defaults you should enforce with SCPs or Config rules:
- S3: block public access at the account level, require encryption on all PUTs, require TLS in transit via bucket policy
aws:SecureTransportcondition. - EBS: turn on default encryption per region. Every new volume is encrypted with your chosen key.
- RDS: require encryption at rest at creation time.
- Secrets Manager and Parameter Store: use CMKs, never plain-text config for anything sensitive.
For cross-account access to encrypted resources, remember that both the resource policy (say, an S3 bucket policy) and the KMS key policy must allow the remote principal. Forgetting the key policy is a top-five cause of "why can't the other account read this bucket" tickets.
Detection: CloudTrail, Config, GuardDuty, and Security Hub
Prevention will fail sometimes. Detection is what keeps a failure from becoming a breach.
CloudTrail is the audit log of every API call in your account. Turn on an organization trail from the management account, deliver logs to the log-archive S3 bucket, and enable log file validation. CloudTrail Lake or Athena queries let you answer questions like "who assumed this role in the last 30 days" or "what IAM changes happened in production last Tuesday". Data events (S3 object-level, Lambda invocation) are extra cost but essential for sensitive buckets and functions.
Config records the configuration state of your resources over time and evaluates them against rules. AWS supplies hundreds of managed rules (s3-bucket-public-read-prohibited, iam-user-mfa-enabled, encrypted-volumes, rds-storage-encrypted) and you can write custom rules in Lambda. Config is how you answer "was this resource compliant on the day of the incident" and how you catch drift from your intended baseline. Enable it in every account, aggregate to the audit account, and route findings to Security Hub.
GuardDuty is threat detection. It ingests CloudTrail, VPC flow logs, DNS logs, EKS audit logs, and (with additional features) S3 data events, RDS login activity, EBS malware scans, and Lambda network activity. It uses ML and threat intelligence to flag things like credential exfiltration to a Tor exit node, cryptomining behavior on an EC2 instance, or an IAM user calling APIs from an unusual geography. Turn it on organization-wide with the audit account as delegated administrator. Findings are noisy at first; tune with suppression rules but be careful not to suppress real signal.
Security Hub aggregates findings from GuardDuty, Config, Inspector, Macie, IAM Access Analyzer, and third-party tools. It also runs standards checks (AWS Foundational Security Best Practices, CIS AWS Foundations, PCI DSS) that produce their own findings. Delegate administration to the audit account so you have one pane of glass across the org. Wire high-severity findings to Slack or PagerDuty via EventBridge, and lower-severity to a ticket queue that your team actually grooms.
IAM Access Analyzer is a specific tool worth calling out. It analyzes resource policies (S3 buckets, KMS keys, IAM roles, Lambda functions, SQS queues, Secrets Manager secrets) and flags any that grant access to principals outside your organization or account. It also generates least-privilege policies from CloudTrail history, which is a genuinely useful way to tighten over-broad policies.
For teams building broader observability practices, the same discipline you would apply to metrics and dashboards, such as those covered in the Prometheus fundamentals walkthrough, applies to security findings: signal versus noise, alert fatigue, and runbook-driven response.
Human access: IAM Identity Center and permission sets
For humans, IAM Identity Center (successor to AWS SSO) is the default. It gives you a single sign-in portal, a permission set model that maps to IAM roles in each account, and integration with any SAML or SCIM identity provider.
Permission sets are collections of policies (AWS-managed, customer-managed by name, and inline) that Identity Center provisions as IAM roles in target accounts. You assign a group from your IdP to a permission set on a specific account or OU. When a user in that group signs in and picks that account, Identity Center creates a temporary session with the corresponding role.
Design permission sets around job function, not around individual accounts. PlatformAdmin, DeveloperReadWrite, DeveloperReadOnly, SecurityAudit, Billing, BreakGlass. Then assign groups to permission sets across the accounts they need. A backend developer might have DeveloperReadWrite on all Non-Prod accounts and DeveloperReadOnly on Prod.
Session duration is a lever. Default is one hour, which is annoying for engineers doing long-running work. Extending to eight or twelve hours is fine for lower-risk permission sets, but keep production admin sessions short, one to two hours, so a stolen laptop with an active session has limited exposure. Require MFA on the IdP side, not just on AWS side, and prefer WebAuthn or hardware keys over TOTP where you can.
Break-glass access needs its own design. Create one or two accounts with IAM users, hardware MFA devices stored in a physical safe, and passwords printed and sealed. These are used only when Identity Center or the IdP is down. Log every use to a channel that pages the security team. Test the process every quarter, because a break-glass procedure you have not rehearsed is not a procedure, it is a wish.
Common IAM mistakes and how to avoid them
Patterns I have seen repeatedly, in rough order of frequency:
Wildcards everywhere. "Action": "*" and "Resource": "*" in policies that were meant to be temporary and never got tightened. Fix: use IAM Access Analyzer's policy generation from CloudTrail history to produce a scoped policy from actual usage, then diff it against what is deployed.
Long-lived access keys on developer laptops. Someone runs aws configure once and the keys live in ~/.aws/credentials forever. Fix: enforce Identity Center for all human access, and use SCPs to deny iam:CreateAccessKey in workload accounts.
Trust policies that trust * or an over-broad principal. A role's trust policy says "Principal": {"AWS": "*"} because someone was debugging cross-account access and never fixed it. Fix: Access Analyzer flags external access explicitly. Also require code review on every trust policy change.
Confused deputy in cross-account roles. You assume a role in a customer's account, but the trust policy doesn't check sts:ExternalId, so any other customer of your SaaS could also assume that role by guessing the ARN. Fix: always require an ExternalId on third-party trust policies, and treat it as a shared secret unique per customer.
Assuming CloudFormation or Terraform state files are safe. State files contain secrets, ARNs, and enough architectural detail to plan an attack. Fix: encrypt the state bucket with a CMK, restrict access to the CI role only, and never commit state to Git.
Overloaded roles. One role does deploys, runs the app, and reads secrets, because someone was in a hurry. Fix: separate deploy-time roles (used by CI) from runtime roles (used by the workload). The deploy role can create resources but does not need to read data; the runtime role reads data but does not need to create infrastructure.
Not scoping resource policies. An S3 bucket policy grants access to an entire account's principals when it only needed to grant to one role. Fix: use aws:PrincipalArn or aws:PrincipalTag conditions to narrow, and let Access Analyzer catch the over-broad ones.
Ignoring service-linked roles. These are roles AWS services create to act on your behalf. They are usually fine, but they show up in audit reports and confuse people. Understand which services created which service-linked roles in your account so you can explain them.
Forgetting the root user. Root has an email, a password, and possibly access keys from years ago. Fix: rotate the root password to something long and stored in a password manager the security team owns, delete all root access keys, enable hardware MFA on root, and never use root except for the specific actions that require it (closing an account, changing support plan, etc.). SCPs cannot restrict root in the management account, which is another reason that account should have nothing else in it.
If you are building foundational skills across these areas, the Refonte cloud engineer program walks through IAM design, landing zones, and incident response with hands-on labs against real AWS environments, so you can practice these patterns before you own them in production.
Blast radius design: containment as a first-class concern
You cannot prevent every incident. You can decide, in advance, how large any single incident is allowed to be. That decision is blast radius design, and it shows up at several layers.
Account-level blast radius. One workload per account for anything sensitive. If a Lambda function is compromised, the attacker's initial context is limited to that account's IAM, that account's data, and that account's network. Cross-account trust is explicit, minimal, and logged.
Network-level blast radius. VPCs are per-account. Security groups are stateful and default-deny inbound. Network ACLs add a subnet-level layer. Egress is controlled through NAT gateways, VPC endpoints, or a shared egress VPC with a proxy. Do not let workloads talk to arbitrary internet endpoints; if a compromised container tries to exfil to a random domain, the network should stop it.
IAM-level blast radius. Each workload role has permissions only for the specific resources it uses, and those resources are usually scoped by name prefix or tag. Permission boundaries cap the maximum. Cross-account role assumption is rare and always uses ExternalId or condition keys.
Data-level blast radius. Encryption keys are scoped. Production data is in a CMK that only production workload roles can decrypt. Backups are in a different key, ideally in a different account. S3 buckets have block-public-access at the account level and object ownership set to bucket-owner-enforced so you cannot accidentally create objects the bucket owner cannot manage.
Time-level blast radius. Credentials expire. Sessions expire. Access keys, if they must exist, rotate. A stolen credential from six months ago should be useless today.
The reliability and cost pillars have blast radius considerations too. A runaway Lambda in a shared account can drain a budget just as effectively as it can leak data, which is why cost optimization on AWS belongs in the same conversation as security. Both are consequences of unbounded workloads in shared environments.
Secrets management: Secrets Manager, Parameter Store, and never in Git
Application secrets (database passwords, third-party API keys, signing keys) belong in Secrets Manager or SSM Parameter Store, never in environment variables baked into an AMI or a container image, never in a Git repo, and never in a CloudFormation template as plain text.
Secrets Manager has native rotation for RDS, Redshift, DocumentDB, and generic secrets via Lambda. It costs $0.40/secret/month, which sounds trivial and adds up if you have thousands of secrets. Parameter Store is free for standard parameters and $0.05/10k requests for advanced parameters. For most cases, Parameter Store with SecureString type and a CMK is enough.
Applications retrieve secrets at startup or lazily on first use, cache them briefly (with jitter to avoid stampeding herd on rotation), and re-fetch on auth failure. Do not embed the secret in the process forever, because you want rotation to actually take effect within a reasonable window.
Access to secrets is via IAM, and the resource policy on the secret should list only the specific runtime roles that need it. Audit access via CloudTrail; unusual patterns like a secret being read by a role that has not read it before are worth an alert.
For build-time secrets (Docker Hub credentials, npm tokens, signing keys used by CI), use OIDC federation from your CI to a role that reads Secrets Manager, rather than storing the actual secret in CI's own secret store. This centralizes secret custody and rotation.
Incident response on AWS: preparation and playbooks
When something goes wrong at 2 a.m., your ability to respond depends on what you set up months ago.
Preparation checklist:
- CloudTrail organization trail is on, delivering to the log-archive account, with log file validation enabled.
- GuardDuty is on in every account and every region, even regions you do not use (attackers love unused regions).
- Config is on with the AWS Foundational Security Best Practices conformance pack.
- Security Hub is aggregating findings in the audit account, with EventBridge rules routing critical findings to PagerDuty.
- Break-glass roles exist in the audit account with permissions to assume read-only roles in any workload account for investigation.
- Playbooks exist for common scenarios: leaked access key, compromised EC2 instance, unauthorized IAM role created, unusual GuardDuty finding in production.
A leaked access key playbook, for example: (1) disable the key immediately via IAM, (2) find the user or role it belonged to and disable it too, (3) pull CloudTrail for the last 90 days filtered by that principal, (4) identify what the key touched (which regions, which services, which resources), (5) rotate any secrets that principal could have read, (6) determine root cause of the leak (Git commit? laptop? shared with vendor?), (7) write up the timeline, (8) file a postmortem with concrete preventions.
Forensics on AWS is easier if you plan for it. Snapshot the EBS volumes of any compromised instance before terminating. Preserve the instance in a stopped state in an isolated security group if possible. Copy VPC flow logs for the relevant time window to a dedicated forensics bucket. Preserve the IAM state (attached policies, session tokens issued) before rotating.
The best incident response is one you have practiced. Run a tabletop exercise every quarter with a fictional scenario. Time how long it takes to answer "what did this key access." If that takes more than an hour, invest in tooling.
Compliance frameworks and how AWS features map to them
If you are subject to SOC 2, ISO 27001, HIPAA, PCI DSS, FedRAMP, or a similar framework, most of the controls map directly to AWS features you should already have on.
- Access control (SOC 2 CC6.1, ISO A.9): IAM Identity Center, MFA, permission boundaries, SCPs, quarterly access reviews.
- Encryption (SOC 2 CC6.7, PCI 3.4/3.5, HIPAA 164.312(a)(2)(iv)): KMS with CMKs, default encryption on S3/EBS/RDS, TLS in transit enforced via bucket policies and load balancer configs.
- Logging and monitoring (SOC 2 CC7.2, ISO A.12.4): CloudTrail organization trail, VPC flow logs, Config, GuardDuty, Security Hub, log retention per policy.
- Change management (SOC 2 CC8.1): infrastructure as code with peer review, deploy-only-from-CI SCPs, CloudTrail evidence of every change.
- Vulnerability management (SOC 2 CC7.1, PCI 6.1): Inspector for EC2 and ECR, patching automation via SSM Patch Manager, container base image scanning in ECR.
- Incident response (SOC 2 CC7.3, ISO A.16): documented playbooks, tabletop exercises, GuardDuty and Security Hub as detection layers.
- Data classification and DLP: Macie for S3, tagging strategy that identifies sensitive data, KMS key scoping per classification.
Auditors will ask for evidence, not opinions. Every control needs a report you can produce on demand. Security Hub's compliance standards produce most of them; supplement with Config conformance pack reports and custom Athena queries against CloudTrail.
If compliance is a driver for your certification plans, look at how the cloud certifications overview sequences the AWS Security Specialty against the Solutions Architect path, because the two overlap heavily on identity and encryption content.
A step-by-step hardening plan for an existing account
If you have inherited a single-account AWS environment that has grown organically, here is an order of operations that works.
- Enable CloudTrail across all regions if it is not already, and deliver logs to a bucket with versioning and MFA delete. You need visibility before you can safely change anything.
- Enable GuardDuty in every region. Let it run for two weeks and read the findings; they will tell you where your worst problems are.
- Enable Config with the AWS Foundational Security Best Practices conformance pack. Do not try to fix everything at once; triage by severity.
- Inventory IAM users and access keys. Delete users nobody recognizes. Rotate keys older than 90 days. Add MFA to every remaining user. This alone closes 80% of easy-attack paths.
- Set up IAM Identity Center and migrate humans off IAM users. Start with the security and platform teams, then engineering, then everyone else.
- Move to AWS Organizations if you are still single-account. Create a management account, move the existing account under it, and set up log-archive and audit accounts.
- Introduce SCPs starting with deny-in-unused-regions and deny-root-usage. These are safe and high-value.
- Add permission boundaries for developer roles that create IAM resources.
- Split environments into separate accounts. This is the biggest lift and takes months, but it is the single most effective containment move you can make.
- Turn on Security Hub with organization aggregation once you have multiple accounts, and start driving down findings.
- Enable Access Analyzer and review external access findings weekly until the list is empty and stays empty.
- Practice incident response with a tabletop exercise involving a compromised access key.
Each step is meaningful on its own. You do not need to complete step 12 to see benefit from step 1.
How AWS security compares to other clouds
If you are working across clouds or evaluating a move, the identity primitives look similar in shape but differ in the details. Azure's RBAC uses role definitions and role assignments at management group, subscription, resource group, and resource scopes, and Entra ID is the identity plane instead of IAM Identity Center. GCP uses IAM roles bound to identities at project, folder, and organization scope, and organization policies play a similar role to SCPs.
The trickier differences are in the assumptions. AWS defaults to explicit, granular policies with a lot of surface area; Azure defaults to broader role definitions with less granularity but simpler mental models; GCP splits identity between service accounts (which are also resources) and human identities in more explicit ways. Each has strengths.
The Azure versus AWS comparison covers the identity model differences in more depth, along with networking and compute. If your team is polyglot across clouds, understanding these differences prevents copying AWS-shaped solutions into Azure where they will not fit.
Serverless and container-specific IAM patterns
Serverless workloads have their own IAM idioms. Every Lambda function has an execution role, and the temptation is to reuse one broad role across many functions. Resist. Each function should have a role scoped to what that function actually does. If the function only reads from one DynamoDB table and writes to one SNS topic, the role should say exactly that.
API Gateway has its own authorization layer (IAM auth, Lambda authorizers, Cognito authorizers) that sits in front of the Lambda's own permissions. Design the two layers together: the authorizer decides if the caller can reach the function, the IAM policy decides what the function can do on their behalf.
For containers, ECS tasks have both a task execution role (used by the ECS agent to pull images and write logs) and a task role (used by the container code itself). Keep them separate. EKS pods should use IAM roles for service accounts (IRSA) or EKS Pod Identity, never node-level roles, so that a compromised pod does not inherit the node's permissions.
The serverless architecture guide walks through the Lambda, API Gateway, and Step Functions permission model in more depth, and much of it applies to any event-driven design.
Automating security: policy-as-code and continuous compliance
Every configuration in this article should be expressed as code and applied by a pipeline, not clicked in the console. Terraform, CDK, Pulumi, and CloudFormation all work. Pick one and be consistent.
Policy-as-code goes further: tools like cfn-nag, checkov, tfsec, and terrascan scan your IaC for known bad patterns before deploy. AWS's own cfn-guard lets you write custom rules in a domain-specific language. Wire these into pull request checks so a wildcard IAM policy or an unencrypted S3 bucket cannot merge to main.
At runtime, Config rules and Security Hub controls catch drift. When a rule fires, either auto-remediate (Config has native remediation actions via SSM documents) or open a ticket. Auto-remediation is powerful but risky; start with high-confidence remediations (removing public access from an S3 bucket, disabling a leaked access key) and expand carefully.
The broader DevOps discipline of shipping fast without breaking things applies directly. Trends in that space, including policy-as-code tooling and platform engineering practices, are covered in the DevOps trends and tooling guide if you want context on where the industry is heading.
Explore the cloud silo
For related pillars on the same topic cluster, work through:
- Refonte cloud learning hub for the full topic map.
- AWS fundamentals if you need the platform primer before diving into IAM.
- AWS Well-Architected framework for how security fits alongside reliability, performance, cost, and operational excellence.
- AWS cost optimization because security controls have cost implications and vice versa.
- Serverless on AWS for Lambda, API Gateway, and event-driven IAM patterns.
- Azure vs AWS for cross-cloud identity comparisons.
- Cloud certifications for a path through the AWS Security Specialty and adjacent credentials.
FAQ
Q: Should I use IAM users at all in a new AWS account? A: Almost never. Use IAM Identity Center for humans and IAM roles for workloads. The only legitimate IAM users are for a small number of legacy integrations that cannot federate, and for break-glass access when your identity provider is down. Even then, enforce MFA and rotate keys.
Q: What is the difference between an SCP and an IAM policy? A: An SCP is a guardrail at the organization or OU level that caps what any principal in an affected account can do. It cannot grant permissions, only restrict. An IAM policy grants permissions to a specific principal or resource. An action is allowed only if both the SCP and the IAM policy allow it.
Q: How many AWS accounts should my company have? A: More than one, and probably more than five. At minimum, separate management, log archive, audit, and workload accounts. For each workload, separate production from non-production. Beyond that, split by team or product as your org grows. Account creation should be cheap and automated.
Q: Do I need customer-managed KMS keys or are AWS-managed keys enough? A: AWS-managed keys are fine for many cases. Use customer-managed keys when you need to control the key policy (who can decrypt), rotate on your own schedule, share across accounts, or satisfy a compliance requirement that mandates customer key control. The cost is $1/month/key plus API calls.
Q: How do I stop developers from creating overly broad IAM roles?
A: Combine three controls. First, require a permission boundary on any role they create, enforced by a condition in the IAM policy that lets them create roles. Second, require peer review on IaC that defines IAM policies. Third, run Access Analyzer and static analysis (checkov, tfsec) in CI to catch wildcard actions and resources before merge.
Q: What is the fastest way to reduce blast radius in an existing single account? A: In order: turn on CloudTrail and GuardDuty, add MFA to every user, add SCPs (deny unused regions, deny root usage) after moving to Organizations, and start migrating workloads out into per-environment accounts. The account split is the biggest single win but takes months. The other steps take days.
Q: How often should I rotate access keys and secrets? A: If you have IAM user access keys at all, rotate every 90 days at most. For secrets in Secrets Manager, rotate database credentials every 30 to 90 days depending on sensitivity, and API keys according to the vendor's guidance. Rotation only matters if the process actually works, so automate it and monitor for rotation failures.
Q: Does turning on GuardDuty, Config, and Security Hub cost a lot? A: In small accounts, tens of dollars per month per service. In large multi-account environments, it can reach thousands. GuardDuty costs scale with CloudTrail events, VPC flow logs, and DNS queries. Config costs scale with configuration items recorded. Security Hub costs scale with findings ingested. Budget for it as a fixed cost of doing business; the alternative (undetected breach) is much more expensive.
