AWS Fundamentals: Regions, VPC, IAM, and the Core Services You Actually Use
Amazon Web Services (AWS) can seem overwhelming until you grasp its core concepts. The mental model for AWS starts with understanding that an AWS account spans multiple regions, each containing Availability Zones (AZs), which are isolated data centers. You interact with AWS resources within these zones, and you secure them using Identity and Access Management (IAM). AWS operates on a shared responsibility model where the cloud provider manages the infrastructure and you manage your data and configurations. In this guide, you’ll learn how regions, AZs, VPCs, and IAM fit together, and get concise overviews of the core services most users need: EC2, S3, RDS, Lambda, CloudFront, Route 53, CloudWatch, SNS, and SQS. As you explore these components, you’ll build a solid foundation before diving into certifications or advanced architectures. For ongoing learning, Refonte offers resources across the Cloud category.
AWS Global Infrastructure: Regions and Availability Zones
AWS is a globally distributed cloud. Each AWS Region is a separate geographic area (like us-east-1 or eu-west-1) that contains multiple Availability Zones (AZs). A region might consist of 3 or more AZs, each AZ being one or more physical data centers with independent power, networking, and cooling (aws.amazon.com) (aws.amazon.com). This design allows you to build highly available applications by distributing resources across AZs. For example, you can launch EC2 instances in two different AZs so that a failure in one zone (due to power outage or natural disaster) won’t bring down your workload in the other AZ. AWS itself emphasizes that each AZ is interconnected with high-bandwidth, encrypted networks, so data replication and load balancing across AZs is fast and secure (aws.amazon.com) (aws.amazon.com).
When you create AWS resources, you pick a region (and optionally an AZ). Your selection affects latency, data residency, and pricing. Customers often choose regions based on proximity to users (lower latency), regulatory compliance (data must reside in certain regions), and availability of specific services (not all services launch everywhere at once). For instance, Asia-Pacific regions might serve Asia-Pacific users faster, while Europe (Frankfurt) keeps data in the EU. AWS is continually adding regions worldwide - as of 2026, there are 39 regions and counting (aws.amazon.com). Each region is independent: resources in us-east-1 do not automatically replicate to us-west-2, for example. If you need disaster recovery across continents, you’ll explicitly replicate data or set up cross-region services like multi-region databases.
Within a region, AZs are your fault-tolerance units. A key rule is to avoid single points of failure: spread your resources across two or more AZs for critical systems. AWS’ mental model: Region > AZ > Subnet/VPC > Resource. A Region is like a city; an AZ is like a neighborhood block. All AWS accounts and resources live in a specific region’s set of AZs (though some global services like IAM work across regions). Understanding this layout is fundamental. For example, AWS lists core services like EC2, S3, RDS, and IAM as available in every region launch (aws.amazon.com), which means when you create those in a region, they start operating immediately across that region’s AZs.
- Choosing a region: Consider compliance (e.g. GDPR requires EU region), latency to end users, cost differences, and service availability. Some new services or features appear first in the US.
- High availability: Deploy in at least two AZs. For example, run a load balancer in front of two EC2 instances in different AZs.
- Global expansion: Use services like Amazon Route 53 (DNS routing) and CloudFront (global CDN) to deliver content worldwide, even if backend systems live in limited regions.
By mastering regions and AZs, you ground your AWS architecture. In short: regions isolate failure domains, AZs provide redundancy, and the AWS global footprint is vast and continually expanding.
Shared Responsibility Model: Security “Of” vs “In” the Cloud
One of the first AWS fundamentals is the shared responsibility model for security and compliance (aws.amazon.com). AWS is responsible for the security of the cloud: the physical data centers, the hardware, the networking, and the hypervisor that powers your virtual machines. Your responsibility is security in the cloud: your data, your applications, your OS on EC2 instances, your configuration, and encryption. AWS clearly illustrates that AWS handles the foundational layers (infrastructure, host OS, virtualization) while you handle everything above (data, accounts, credentials, firewall rules) (aws.amazon.com).
For example: - When you use S3 or RDS, AWS ensures the service itself is patched and secured. You are responsible for access policies (who can list or get data in your S3 bucket). - When you run EC2 instances, you must maintain the guest OS and handle security patches, just as you would on a server in your own data center. - Network security (security groups and NACLs) is configured by you. AWS provides the network isolation but you define which ports are open to the world.
This model prevents you from being overwhelmed by compliance and infrastructure tasks. However, it also means you must proactively secure your resources. Always apply least privilege in IAM policies (grant only necessary permissions), enable encryption for data at rest and in transit, and keep your software updated. For detailed strategies, see Refonte’s AWS Security & IAM guide for best practices. In practice, this model means combining AWS’s built-in protections (like VPC isolation and AWS-managed patching in managed services) with your own security processes (like using IAM roles instead of long-term keys).
Major compliance standards (PCI DSS, HIPAA, etc.) note that AWS inherits a lot of burden but customers still need to configure things correctly. Think of AWS as your server host: it locks the data center door, but you lock the door to your room inside.
Virtual Private Cloud (VPC) and Networking
Amazon VPC is the cornerstone of AWS networking. A VPC is a logically isolated virtual network that you define (docs.aws.amazon.com). It “closely resembles a traditional network in a data center,” but with the scalability and flexibility of the cloud (docs.aws.amazon.com). When you create an AWS account, AWS automatically provides a default VPC in each region. You can launch resources into this default VPC or create your own VPCs with custom IP ranges.
Key VPC concepts:
- CIDR block: Each VPC has an IPV4 network range (e.g., 10.0.0.0/16). You subdivide this into subnets (e.g., 10.0.1.0/24 and 10.0.2.0/24) across AZs.
- Subnets: Subnets are segments of the VPC range in a specific AZ. You typically create at least one public subnet (with internet access via an Internet Gateway) and one private subnet (no direct internet) per AZ for redundancy.
- Route Tables: Define the traffic routing. A public subnet’s route table might send 0.0.0.0/0 (all Internet traffic) to an Internet Gateway. A private subnet’s route table might send Internet traffic to a NAT Gateway (so instances can reach out to the internet but are not directly reachable from outside).
- Internet Gateway (IGW): Attached to the VPC to allow ingress/egress to the internet for subnets with the right route.
- NAT Gateway: A managed AWS service that lets instances in private subnets access the internet (for updates, etc.) without exposing them to inbound internet connections.
- Security Groups (SGs): Act as virtual firewalls at the instance level. For example, you can create a security group allowing port 80 traffic only from the internet or from specific IP ranges.
- Network ACLs: Stateless firewalls at the subnet level. Less commonly edited than security groups, often kept open or very restrictive defaults.
- VPC Peering and Transit Gateway: For connecting VPCs (even across accounts or regions) privately.
Creating a VPC with CLI (example in Terraform pseudocode, for clarity):
resource "aws_vpc" "main" {
cidr_block = "10.0.0.0/16"
enable_dns_hostnames = true
}
resource "aws_subnet" "public" {
vpc_id = aws_vpc.main.id
cidr_block = "10.0.1.0/24"
availability_zone = "us-east-1a"
}
resource "aws_internet_gateway" "igw" {
vpc_id = aws_vpc.main.id
}
resource "aws_route_table" "public" {
vpc_id = aws_vpc.main.id
route { cidr_block = "0.0.0.0/0"; gateway_id = aws_internet_gateway.igw.id }
}
resource "aws_route_table_association" "pub-assoc" {
subnet_id = aws_subnet.public.id
route_table_id = aws_route_table.public.id
}
This example creates a new VPC with one public subnet and an Internet Gateway with a default route. AWS Console or CLI can do the same via aws ec2 create-vpc, create-subnet, create-route-table, etc. Always plan your IP ranges to avoid overlaps if you need to connect VPCs or to an on-premise network.
A practical example: Suppose you have a web application. You might launch your web server EC2 instances in a public subnet (so they have public IPs and receive traffic from the internet via a Load Balancer). These servers can then connect to a database instance in a private subnet, which has no internet access. The database can reach AWS update servers via a NAT Gateway, but outsiders cannot directly access it. All traffic rules are enforced by security groups and NACLs.
Networking in AWS can get complex, with advanced topics like VPC Endpoints (to privately connect to AWS services without using the internet), VPN/Direct Connect to on-prem networks, and IPv6 support. For many, the core is: VPCs are your private networks, subnets and route tables determine traffic flow, and SGs/NACLs lock things down. Properly using VPC and networking is critical, so new AWS users often spend significant time on it.
Example step-by-step: To set up a simple 3-tier network: 1. Create a VPC with range 10.0.0.0/16.
2. Create two public subnets (10.0.1.0/24 in AZ1, 10.0.2.0/24 in AZ2).
3. Create two private subnets (10.0.3.0/24 AZ1, 10.0.4.0/24 AZ2).
4. Attach an Internet Gateway to the VPC.
5. Create a public route table and route 0.0.0.0/0 to the Internet Gateway; associate public subnets with it.
6. Create a NAT Gateway in the public subnet(s).
7. Create a private route table, route 0.0.0.0/0 to the NAT Gateway; associate private subnets.
8. Launch a public EC2 in one public subnet, set its SG to allow HTTP, SSH (for admin).
9. Launch a private EC2 (or RDS) in private subnet, allow only database traffic from the public instance’s SG.
This design ensures your backend is isolated while still reachable by your front-end servers.
Identity and Access Management (IAM)
AWS IAM (Identity and Access Management) is how you control who (or what) can do what in your cloud. IAM allows you to create principals (users, groups of users, and roles) and policies (permissions) to specify allowed actions on resources. Every action in AWS is performed by a principal. You should avoid using the root account for daily tasks; instead, create an IAM user or role with limited permissions.
Key IAM elements:
- Users: Individual identities for people or applications. For example, create an IAM user for Alice, another for Bob.
- Groups: Collections of IAM users under the same permissions. For example, a “Developers” group that has permission to manage certain resources.
- Roles: Non-human identities that AWS services or external federated users can assume. For example, you might assign an IAM Role to an EC2 instance so that the instance can access S3 without storing credentials on the instance.
- Policies: JSON documents that define permissions (action and resource). They attach to users, groups, and roles. AWS has many managed policies (predefined), and you can write custom policies. For example, an IAM policy might allow s3:GetObject on a specific S3 bucket.
Example IAM policy (JSON) granting list and read access to one S3 bucket:
{
"Version": "2012-10-17",
"Statement": [{
"Effect": "Allow",
"Action": ["s3:ListBucket"],
"Resource": ["arn:aws:s3:::example-bucket"]
},{
"Effect": "Allow",
"Action": ["s3:GetObject"],
"Resource": ["arn:aws:s3:::example-bucket/*"]
}]
}
This policy could be attached to a user or role. When the principal tries to list or get objects from example-bucket, these rules allow it (assuming no other policy denies it).
Best practices: - Enable MFA (Multi-Factor Authentication) for privileged accounts. - Follow least privilege: only grant minimal permissions. - Use roles instead of embedding access keys in code or servers. For example, assign an EC2 Role with an S3 access policy to the EC2. Then applications on that EC2 automatically get credentials from the Role. - Rotate keys regularly if you use access keys (especially for programmatic access). - When integrating with other accounts, use Cross-account roles and trust relationships. - Keep sensitive data out of policies (no passwords or secrets in IAM, use AWS Secrets Manager or SSM Parameter Store for secrets).
IAM has a learning curve (policies can be verbose), but it’s central to AWS security. For detailed security strategies, see the AWS Security & IAM guide on our site. In short, think of IAM as the lock-and-key system for AWS: every user or service needs the right key (role or policy) to unlock (access) a resource.
Compute Services: Amazon EC2
Amazon EC2 (Elastic Compute Cloud) is AWS’s core virtual server service. EC2 provides resizable compute capacity (virtual machines) in the cloud. You can choose instance types (CPU, RAM, networking) to match your workload. For example, a t3.micro has 2 vCPUs and 1 GiB of memory (good for small tasks or learning), while an m5.large has 2 vCPUs and 8 GiB, and c5.4xlarge has 16 vCPUs and 32 GiB (for CPU-intensive tasks).
To launch an EC2 instance, you typically: 1. Select an AMI (Amazon Machine Image), which is a template for the OS and software (e.g., Amazon Linux, Ubuntu, Windows). 2. Choose an Instance Type based on needed compute and memory. 3. Select a Region and AZ. (An AMI may be region-specific, but common ones exist everywhere.) 4. Pick a VPC/subnet (public or private). 5. Assign a security group for network access. 6. (Optional) Attach storage (EBS volumes) and key pairs for SSH. 7. Launch.
Here is an example AWS CLI command to launch a t3.micro in us-east-1:
aws ec2 run-instances --image-id ami-0abcdef1234abcd12 --instance-type t3.micro --count 1 --security-group-ids sg-0123456789abcdef0 --subnet-id subnet-0abc1234def --key-name MyKeyPair
This command starts an EC2 instance using the specified AMI. After a minute, the instance is running and has a public IP (if in a public subnet with an IGW). You could then SSH into it using the MyKeyPair.
EC2 scaling options: - Auto Scaling: Automatically launch or terminate instances based on load (via Auto Scaling Groups and metrics). - Elastic Load Balancing (ELB): Distributes traffic across instances. - Spot Instances: Bid on spare capacity for a very low price (good for fault-tolerant or non-urgent jobs). - Reserved Instances/Savings Plans: Commit to 1-3 years to save on costs for steady workloads.
EC2 is highly flexible but requires management of the OS. If you place a web server on an EC2, you must patch and maintain it (for Linux or Windows). For fully managed containers or serverless alternatives, AWS offers ECS, EKS, or Lambda (see next section). However, EC2 is often the starting point for understanding AWS.
Use cases: - Legacy lift-and-shift: Running an existing application on a cloud server. - Web servers, application servers, batch jobs. - Any workload needing a full OS environment.
Comparison (EC2 vs others): EC2 vs AWS Lambda (serverless) vs on-premises can be summarized as: | Service | Category | Ideal For | |-------------|--------------|-------------------------------------| | Amazon EC2 | IaaS Compute | Traditional VMs, full OS control | | AWS Lambda | Serverless | Event-driven tasks, bursts, microservices | | On-Prem | Private Host | Legacy hardware, sensitive data (if needed) |
In the table, note AWS also has container services (ECS/EKS) that offer more control than Lambda but less overhead than EC2. For fundamental learning, run a few EC2 instances in a VPC and get comfortable with SSH, security groups, and storage attachment.
Serverless Compute: AWS Lambda
AWS Lambda is AWS’s serverless compute service. Instead of managing servers, you define functions (code) that run in response to events. Lambda automatically manages scaling: if hundreds of events come in, Lambda spins up as many instances of your function as needed.
How it works: - You write a function (in Node.js, Python, Go, Java, etc.) and upload it to Lambda. - You set up a trigger: this could be an API Gateway HTTP call, an AWS event (like an S3 object upload), a scheduled CloudWatch Event (cron), or a message on SNS/SQS, among others. - When the event occurs, Lambda runs your function in a micro-VM, passes the event data to the function, and executes your code. - You pay only for the compute time used (in increments of 100ms) plus request count. No charge when no one is invoking your function.
Example Lambda use-case: Generate a thumbnail whenever a user uploads an image to S3.
1. User uploads photo.jpg to an S3 bucket.
2. S3 triggers your Lambda function with event data (bucket name, key).
3. Lambda code retrieves photo.jpg, uses an image library to create a thumbnail.jpg.
4. Lambda uploads thumbnail.jpg to a different S3 bucket or folder.
All without you provisioning or maintaining a server.
Lambda constraints: - Execution time limit (up to 15 minutes, but often used for sub-second tasks). - Limited memory (up to 10 GB) and ephemeral storage (512 MB temp space during execution). - Stateless (for each invocation, assume no previous data in memory or file system).
Lambda excels for: - APIs: Combined with Amazon API Gateway, you can build REST APIs without servers. - Data processing: Respond to S3 or DynamoDB events to transform data. - Automation: Scheduled tasks (e.g., nightly jobs) via CloudWatch Events. - Glue between services: E.g., SNS triggers Lambda to send emails, or SQS triggers Lambda to process a queue item.
Since Lambda functions are short-lived, persistent connections or long-running tasks don’t fit well. For those, EC2 or containers might be better. But as part of AWS fundamentals, remember: lambda = code as a service. It offloads all server management to AWS. If you plan to build modern microservices, learn Lambda and check Refonte’s guide on serverless architectures for deeper coverage.
Object Storage: Amazon S3
Amazon Simple Storage Service (S3) is AWS’s scalable object storage. Think of S3 as a huge, secure, highly durable file system in the cloud. You store objects (files) in buckets. Each object is identified by a key (filename) and has data plus metadata.
Key features of S3: - Scalability: Virtually unlimited capacity; S3 will grow as needed. - Durability: AWS design for 99.999999999% (11 nines) durability. They store copies across multiple facilities. - Tiers: Standard, Infrequent Access, and Glacier for archiving as data ages. - Access control: Bucket policies and ACLs let you share or restrict data. - Versioning: Keep old versions of objects to protect against accidental deletes. - Static website hosting: You can host an entire static website from an S3 bucket by enabling the website feature and uploading HTML/CSS. - Notifications and Events: S3 can trigger Lambda functions or send messages to SNS/SQS on events (e.g., object created or deleted). - Encryption: AWS-managed keys (SSE-S3), KMS-managed keys (SSE-KMS), or client-side encryption.
Example: Uploading a file with AWS CLI:
aws s3 cp report.pdf s3://company-reports/2026/financials.pdf
This command copies report.pdf from your machine to the specified S3 bucket and key.
S3 is key-value object storage, not a traditional filesystem. You don’t log into S3 - you use APIs, SDKs, CLI, or the console. For example, if you have s3://myassets/images/logo.png, you retrieve it with
aws s3 cp s3://myassets/images/logo.png ./logo.png
S3 scales massively - it can serve millions of requests per second if set up with CloudFront in front.
Costs: You pay for storage (per GB) and data transfer out (egress) and API requests (PUT/GET etc). Frequent retrieval costs more, so AWS encourages placing seldom-accessed files in cheaper classes.
Use cases: - Data lakes or backups: Store large amounts of data cost-effectively. - Static content: Host images, videos, or static web assets. - Big data: Store inputs/outputs for analytics pipelines. - As part of workflows: e.g., records from an application, logs archives.
Comparisons: | Feature | S3 (object) | EBS (block) | EFS (file) | |--------------|---------------|------------------|-------------------| | Storage type | Object | Block (attached) | Network file system | | Use case | Backups, static hosting | Root disks for EC2/DB | Shared files across servers | In this table, S3 is unique for scalability and direct internet accessibility (each object has a URL).
Example project: Host a static site. Steps:
1. Create a bucket named example-site.
2. Enable static website hosting (set index.html as the homepage).
3. Upload index.html, styles.css, etc. to the bucket.
4. Configure bucket policy to allow public read on these objects.
5. Now http://example-site.s3-website-us-east-1.amazonaws.com serves your site.
For dynamic content, sometimes S3 is paired with CloudFront (CDN) to accelerate delivery globally. S3 often ties into many AWS services: it’s a common source for backups (RDS can snapshot to S3), source/target for data analytics, and origin for CDNs.
Relational Databases: Amazon RDS
Amazon RDS (Relational Database Service) is AWS’s managed relational database solution. Instead of running MySQL or PostgreSQL on an EC2 yourself, you tell RDS, “I need a MySQL/PostgreSQL/Oracle/SQL Server/MariaDB/Amazon Aurora instance”. AWS then provision and manages a database instance. You still choose instance size, engine, and initial setup, but AWS handles patching, software installation, backups, and (optionally) failover.
To use RDS, you: 1. Pick a database engine (e.g., MySQL 8.0, PostgreSQL 15, etc.). 2. Specify an instance class (e.g., db.t3.micro for dev/test, db.m5.large for production). 3. Define storage (gp3 SSD, amount in GB). 4. Select Multi-AZ deployment if you want standby replication (automatic failover). 5. Configure VPC and subnets (RDS sits in a subnet; often private). 6. Create or provide a DB name, username, and password.
AWS then creates the database instance. You connect via its endpoint (DNS name and port). For example, mydb.abc123xyz.us-east-1.rds.amazonaws.com:3306. RDS reserves an IP internal to your VPC.
RDS features:
- Automated backups: Daily snapshots and point-in-time recoveries.
- Multi-AZ: Creates a synchronous standby in another AZ for failover. If the main AZ goes down, AWS automatically fails over to the standby with a brief outage (usually less than a minute).
- Read Replicas: For read-heavy workloads, you can create read-only replicas in the same region (or even cross-region with Aurora/Postgres).
- Scaling: You can vertically scale (increase instance class) or storage on the fly.
- Managed maintenance: AWS patches the DB engine, though you can control maintenance windows.
- Parameter groups: Tweak database settings, similar to my.cnf.
- Encryption: Use AWS KMS keys to encrypt data at rest.
Example AWS CLI to create a MySQL RDS (simplified):
aws rds create-db-instance \
--db-instance-identifier mydb \
--db-instance-class db.t3.small \
--engine mysql \
--master-username admin \
--master-user-password Secret123 \
--allocated-storage 20
It will be accessible once created. You’d then configure your app to use it.
RDS vs DynamoDB vs self-hosted: RDS is for relational databases with SQL, multi-row transactions, complex queries. If you need auto-scaling of reads and writes without worrying about SQL, consider DynamoDB (NoSQL) - not covered here but vital to know. RDS keeps you in familiar RDBMS land but moves the operational burden to AWS.
Use cases: - Traditional web applications that rely on MySQL/PostgreSQL/Oracle/SQL Server. - Applications needing long transactions or complex joins. - Enterprises migrating existing DBs to cloud with minimal changes.
AWS also offers Amazon Aurora, a special high-performance MySQL/PostgreSQL compatible RDS variant. Aurora auto-scales storage and has a cluster of readers/writers. But Aurora is more advanced; start with standard RDS instances.
A common scenario: Customer runs a blog. As traffic grew, they got tired of managing their own MySQL on EC2 and suffering downtime during patching. They switch to RDS and set up a read replica for reporting queries. Now AWS handles patching and failover, and they only worry about database credentials and schema.
Important note: With RDS, you don’t get shell access to the underlying OS or direct hardware control. It's a service endpoint. But you do manage the data, instance settings, and can run SQL commands as the DB admin user. Make sure you secure it: set the RDS SG to allow access only from your application server group.
Content Delivery Network: Amazon CloudFront
Amazon CloudFront is AWS’s Global Content Delivery Network (CDN). It caches your content at edge locations worldwide (over 300 in 95+ cities) so that users download from a nearby point, improving performance. CloudFront can accelerate delivery of any static or dynamic content: HTML, CSS, images, videos, APIs, etc. Even non-HTTP content (e.g., RTMP, though that’s legacy) can be served.
Here's how CloudFront works in a nutshell:
- You create a distribution in CloudFront. It has one or more origins (the source of content). Typical origin: an S3 bucket or an HTTP server (like an EC2 load balancer or external site).
- You assign a domain name for your distribution, like d123.cloudfront.net. You can also add your own domain (via Alternate Domain Names and certificate in AWS Certificate Manager).
- Geographic areas fetch from their nearest edge. When content is requested by users, CloudFront checks if it has a cached copy in that location. If not, it pulls from the origin and caches it for subsequent requests (based on TTL).
- You pay for data transfer and requests at the edge, which is usually cheaper than direct from the origin, and often costs less overall when factoring lower origin costs and better performance.
Example use-case: You have a website hosted on S3 and use CloudFront:
1. Set your S3 bucket as the CloudFront origin (with an origin path if needed).
2. Configure the distribution behaviors (cache policies, time-to-live, allowed HTTP methods).
3. CloudFront domain (e.g. dexample.cloudfront.net) or custom domain (e.g. static.example.com).
4. When a user accesses static.example.com/logo.png, CloudFront edge in, say, Singapore checks cache; if not present, retrieves logo.png from S3 and caches it.
One key feature is support for HTTPS easily via AWS-managed certificates. You can serve secure content without handling SSL certificates yourself - CloudFront does it.
Comparing with other AWS services: | Aspect | Direct S3/EC2 (origin) | CloudFront CDN | |-------------------|------------------------|------------------| | Latency | Depends on region | Edge (low latency globally) | | Throughput | Often region-limited | High (multiple edges) | | Caching | No | Yes (reduces origin load) | | Cost | Higher egress charges | Lower egress at edge, plus request pricing | CloudFront is almost always a win for static assets or any regionally-shared content. Even for APIs, you can use CloudFront with Lambda@Edge for logic at edge nodes (though that’s advanced).
Security: CloudFront integrates with AWS WAF (firewall) and Shield (DDoS protection). You can restrict S3 buckets so only CloudFront can fetch (via OAC origin access control), keeping your S3 private except through CloudFront.
Example Workflow
Suppose you have an online store. You’d use CloudFront as follows:
- All images, CSS, JS, and even backend API calls go through CloudFront.
- Create an origin for the S3 bucket of images and another origin for your API load balancer.
- Set caching behavior: static files can have a long TTL (e.g., 1 day), API calls maybe no-cache or short.
- Callers use a custom domain, e.g. api.yourstore.com, which is CNAME’d to the CloudFront distribution. Use ACM for the SSL certificate on that domain.
- CloudFront layers: Edge caches, then if needed, regional cache (if using regional edge cache), then origin.
By doing this, customers globally see faster load times and less strain on your origin servers. CloudFront’s pricing is pay-as-you-go. It can get complex (nation-based rates), but overall it's straightforward: data out of CloudFront edges is cheaper than data out directly from origin servers in EC2 or S3.
If you want to learn more about optimizing performance and costs with Cloud architecture, the AWS Well-Architected Framework is a great resource to lean on for best practices.
Domain Name Service: Amazon Route 53
Amazon Route 53 is AWS’s scalable DNS and domain registration service. It translates human-friendly names (like www.example.com) to IP addresses or AWS services. Beyond basic DNS, Route 53 provides advanced routing policies, health checks, and domain name registration (you can buy and manage domain names directly).
Key features:
- Hosted Zones: A collection of DNS records for a domain, e.g. example.com. You create a hosted zone, then add records like A, CNAME, MX, etc.
- Record Types: All standard DNS types are supported. For example:
- A (IPv4 address)
- AAAA (IPv6)
- CNAME (alias)
- MX (mail exchange)
- TXT (text, e.g., for verification)
- NS (name servers)
- Alias (AWS-specific: route to S3 static site, CloudFront, ELB etc at zone apex).
- Routing Policies:
- Simple: straightforward one record one IP.
- Weighted: distribute traffic among endpoints (like AB testing or gradual rollout).
- Latency-based: respond with the lowest-latency AWS region for the requester.
- Geolocation: serve different content based on the user’s location (region, country).
- Failover: active/passive routing. Use health checks to switch to a secondary endpoint if primary fails.
- Multivalue Answer: like simple round-robin but with health checks on multiple A records.
- Health Checks: Route 53 can continuously check if a resource (like a web server) is up. If health checks fail, Route 53 can stop sending traffic there (for failover routes).
- Domain Registration: You can buy and transfer domain names (.com, .org, .net, many countries, etc) directly in Route 53.
Example scenario: You have two AWS regions (us-east-1 and eu-west-1) running instances for your service. You want users to use the nearest region. In Route 53, you create a latency-based routing policy with two A records (one pointing to the ELB DNS of us-east-1 and one to eu-west-1). When a user in Europe resolves api.example.com, Route 53 gives the EU ELB address; a user in the US gets the US ELB address.
Setting up a basic record via CLI (note the --name and --type fields):
aws route53 change-resource-record-sets --hosted-zone-id Z1D633PJN98FT9 --change-batch '{
"Changes": [{
"Action": "UPSERT",
"ResourceRecordSet": {
"Name": "api.example.com",
"Type": "A",
"AliasTarget": {
"HostedZoneId": "Z35SXDOTRQ7X7K", # hosted zone for ELB
"DNSName": "my-elb-123.us-east-1.elb.amazonaws.com",
"EvaluateTargetHealth": false
}
}
}]
}'
This creates or updates an A record using Alias (pointing to an AWS Load Balancer). Alias records are better than CNAME at the zone apex (root domain) because they don't incur costs and can use AWS resources as targets.
Use cases: - Use Route 53 as your primary DNS. It’s reliable and quick globally. - For routing across multiple AWS accounts, you can share hosted zones or duplicate records. - Domain testing: Use weighted records to send 10% of traffic to a new server, 90% to old. If metrics are good, switch to 100%. - Disaster recovery: Pair with health checks. If one region’s app fails its health check, Route 53 can route all users to the backup region.
Route 53 is often paired with CloudFront (you point your domain to the CloudFront distribution's domain) or with Elastic Load Balancers. Together with services like AWS Certificate Manager (for SSL certs) and CloudFront or ALB, you can build secure, global web architectures.
For multi-cloud scenarios (e.g., an app partly on Azure and on AWS), Route 53 still manages the DNS for the domain, but might have multiple A records in different clouds. (We compare multi-cloud architecture in our Azure vs AWS guide if interested.)
Monitoring and Logging: Amazon CloudWatch
Amazon CloudWatch is the monitoring and observability service in AWS. Think of it as the central place for your metrics, logs, and alarms. It gives insight into performance and operational health across AWS services.
Key CloudWatch features:
- Metrics: AWS services emit metrics (like CPUUtilization for EC2, or Latency for your Elastic Load Balancer) that CloudWatch collects. By default, basic EC2 metrics (CPU, disk, network) are 5-min granularity, and you can enable detailed monitoring (1-min).
- Custom metrics: You can send your own metrics to CloudWatch (from your apps, using the AWS SDK or aws cloudwatch put-metric-data). For example, track application-level stats like “orders per minute”.
- Dashboards: Create visual dashboards showing graphs of metrics. For example, a web admin might see a chart of requests per second and error rate.
- Logs: CloudWatch Logs can collect logs from:
- EC2 (via the CloudWatch agent or a logging driver).
- Lambda (each function’s stdout and stderr go to CloudWatch Logs).
- Elastic Beanstalk, Amazon RDS logs (let AWS push them).
- On-prem or any host (via the agent).
You configure Log Groups and Log Streams, then you can search logs or set retention policies.
- Alarms: You set a CloudWatch Alarm on a metric (e.g., “CPUUtilization > 80% for 5 minutes”). The alarm can trigger actions: send an SNS notification, scale up an Auto Scaling group, or trigger a Lambda.
- Events (EventBridge): CloudWatch Events (now Amazon EventBridge) can respond to changes in AWS (like an EC2 state change, or a new AMI creation) on a schedule or a pattern, invoking targets (Lambda, SNS, SQS etc).
Examples:
- Create a CPU alarm: If an application load peaks, you might want an email. aws cloudwatch put-metric-alarm --alarm-name HighCPU --metric-name CPUUtilization --namespace AWS/EC2 --statistic Average --period 300 --threshold 75 --comparison-operator GreaterThanThreshold --dimensions Name=InstanceId,Value=i-0123456789abcdef0 --evaluation-periods 2 --alarm-actions arn:aws:sns:us-east-1:123456789012:MyTopic
This sets an alarm on one EC2 instance’s CPU.
- Stream logs: Install CloudWatch Agent on your EC2 or on-prem server, configure it to send /var/log/syslog to a CloudWatch log group. Then you can filter-log-events via CLI or see them in the console.
Monitoring vs Logging vs Tracing: CloudWatch primarily does metrics & logs. For distributed tracing (e.g., microservices call chains), AWS X-Ray is another service (beyond scope here). But CloudWatch is often first used to get an overall health picture.
Billing and Cost: Light usage of CloudWatch (basic metrics and threshold alarms) is usually free or low cost. Custom metrics and high-cardinality (many unique dimensions) can add costs. For logs, you pay per GB ingested and stored. Best practice: filter/rotate logs to avoid runaway costs (Log Groups let you set retention like 30 days).
Use cases: - Automated scaling: An Auto Scaling Group might scale out when average CPU > 70%. - Container monitoring: ECS tasks publish metrics to CloudWatch. - Insights: Identify memory leaks or errors by searching log text in CloudWatch Logs console. - Alerts: If RDS free storage space is low, alarm to send Slack/Email.
CloudWatch ties into almost every AWS service, so it’s a fundamental part of the AWS mental model. For centralized logging across environments (EKS clusters, on-prem, etc.), people also consider tools like Prometheus/Grafana or ELK, but AWS’s native stack is CloudWatch. (If you enjoy open-source monitoring, check our free Prometheus tutorial for alternatives or complementary tools.)
Messaging and Notifications: Amazon SNS and SQS
Decoupling components of your application is easier with AWS messaging services. The two main ones are SNS (Simple Notification Service) and SQS (Simple Queue Service). They serve different but complementary purposes:
- SNS (Simple Notification Service) is a publish/subscribe (pub-sub) system. You publish messages to a topic, and SNS pushes those messages to all topic subscribers. Subscribers can be HTTP/S endpoints, email addresses, mobile SMS, AWS Lambda functions, SQS queues, or even other AWS services. SNS is push-based.
- SQS (Simple Queue Service) is a message queue service. Producers (senders) put messages into a queue. Consumers poll the queue and process the messages. SQS retains messages until a consumer retrieves and deletes them. It’s pull-based.
Here’s a concise comparison:
| Feature | Amazon SNS (Notification) | Amazon SQS (Messaging Queue) |
|---|---|---|
| Pattern | Pub/Sub (push to subscribers) | Queue (pull by consumers) |
| Delivery | Pushed to endpoints (HTTP, email, Lambda, SQS, SMS) | Pulled by consuming clients |
| Ordering/Delivery | Best-effort (some duplication possible) | FIFO available (if enabled) |
| Typical Use-case | Real-time notifications/fan-out | Task queues, decoupling workers |
For example: - Use SNS if multiple subscribers must get updates simultaneously. E.g., an order is placed-send an email to the customer, notify the warehouse system via HTTP, and message a monitoring service. - Use SQS for task processing. E.g., a web app receives orders and places each order message into an SQS queue. A fleet of worker processes polls the queue, processes orders one by one, then deletes them. SQS smooths out spikes: if 100 orders come at once, all sit in queue and workers gradually pull them.
SNS can even send messages to SQS. That combination gives you a fan-out pattern: one SNS topic to many SQS queues (or Lambda functions). Each SQS subscriber will have its own copy of the message, effectively distributing work to multiple consumers.
Example Usage: - An application generates an alert. It publishes to an SNS topic. SNS sends emails to admins and HTTP-posts to a logging service. - A chat application queues notifications: user messages go to an SQS queue; worker Lambdas read the queue and deliver messages to recipients. - SNS + SQS: A single topic for order events; it has two SQS subscribers - one for inventory processing, one for billing. Each gets the same order message but handles it independently.
Sample Commands: Publish message to SNS:
aws sns publish --topic-arn arn:aws:sns:us-east-1:123456789012:UserSignUpTopic --message "New user signed up with email [email protected]"
Send message to SQS:
aws sqs send-message --queue-url https://sqs.us-east-1.amazonaws.com/123456789012/MyQueue --message-body "orderID=12345"
Receiving from SQS:
aws sqs receive-message --queue-url https://sqs.us-east-1.amazonaws.com/123456789012/MyQueue
(after processing, remember to delete-message to remove it).
Reliability: - SNS is near real-time (sometimes seconds delay, but it's usually fast). - SQS guarantees at-least-once delivery. Messages might occasionally be delivered more than once and out of order (unless using FIFO queues, which guarantee order and exactly-once processing but have limits on throughput).
Costs: Both SNS and SQS are very cheap. SNS charges per published message and per notification delivery; SQS charges per request (send/receive/delete operations) and queue storage. Even a busy application usually incurs minimal cost.
Integration: SNS & SQS connect well with AWS Lambda (Lambda can be an SQS consumer or SNS subscriber) and CloudWatch (you can trigger Lambda on SQS queue length alarms, etc.). They also integrate with AWS IoT for sensor messages, and with third-party like Slack via webhooks HTTP endpoints.
Visual Example: Here’s a simple diagram of an order-fanout:
[ New Order via API ]
|
SNS Topic (subscriber -> HTTP service)
/ \
/ \
SQS1 SQS2 (two worker queues)
When an order arrives, your app publishes to SNS “New order: #123”. SQS1 and SQS2 each receive a copy of the message. Maybe SQS1 workers charge credit cards, SQS2 workers update inventory.
For more on decoupling and asynchronous communication in AWS, exploring these services is key. They tie into design patterns covered by AWS and DevOps. (For example, our DevOps 2026 trends article discusses microservices and messaging as part of modern cloud architectures.)
Best Practices, Architectural Pillars, and Next Steps
Having covered the core pieces - from infrastructure (Regions, VPC, AZs) to security (IAM, Shared Responsibility) to pivotal services (compute, storage, networking, database, monitoring, and messaging) - the next step is to put these fundamentals into practice and align them with best practices. AWS themselves champion the Well-Architected Framework, which outlines five pillars (security, reliability, performance, cost optimization, operational excellence). You should design your workloads to meet these pillars; for example, by using multi-AZ deployments (reliability), encrypting data (security), and using auto-scaling (cost/performance). Our AWS Well-Architected page provides strategies and questions to ensure sound architecture designs.
Security best practices are covered in the AWS Security & IAM guide. Topics like least privilege, rotating keys, and network controls (security groups, AWS WAF) are crucial for protecting your resources beyond the shared responsibility baseline.
Cost optimization is another important aspect. AWS provides tools like Trusted Advisor, Cost Explorer, and services like AWS Budgets. Our AWS Cost Optimization page explains concepts like right-sizing compute, using spot instances, and archival storage to save money.
If you’re learning AWS for a career, consider formal certifications. Understanding fundamentals makes certification study easier. Check out our cloud certifications guide for paths like the AWS Certified Solutions Architect or Developer tracks. Certifications often cover these topics extensively, and knowing how services fit together (the mental model) is critical.
For practical experience, hands-on projects are invaluable. Refonte’s Cloud Engineer Program is a training and internship program where you can work on real AWS projects, guided by mentors. It’s an excellent way to apply what you’ve learned here in a structured environment, building skills employers value.
Finally, keep learning about AWS’s evolving ecosystem. For example, containers (ECS/EKS), machine learning services, and IoT are large areas outside this “top 10” list but are worth exploring as you grow. Also, consider how AWS fits into broader IT - for example, Azure vs AWS comparisons if you ever work in a multi-cloud context, or integrating with monitoring tools like Prometheus.
In summary, you now have the core AWS mental model and understanding of its primary services. The cloud is vast, but AWS fundamentals - regions/AZs, VPC, IAM, and core services like EC2/S3/RDS/etc. - are the building blocks. As you build on this foundation, regularly refer back to these concepts. They’ll reappear in advanced topics like microservices, Big Data on AWS, and DevOps workflows.
FAQ
Q: What is the AWS Shared Responsibility Model?
A: AWS’s shared responsibility model means AWS secures the infrastructure (data centers, hardware, network, hypervisor), while you secure the resources you create (your data, operating system, applications, and configurations). For example, AWS ensures S3 is redundant and patched, but you must set the right access permissions on your buckets (aws.amazon.com).
Q: How do AWS Regions and Availability Zones work?
A: An AWS Region is a geographic location containing multiple Availability Zones (AZs). AZs are physically separate data centers within a region (aws.amazon.com) (aws.amazon.com). You deploy resources in a region and can spread them across AZs for fault tolerance. For instance, running a database in two AZs protects against a single AZ outage.
Q: What is Amazon VPC and why use it?
A: Amazon VPC (Virtual Private Cloud) is your private network space in AWS (docs.aws.amazon.com). It lets you define subnets, IP ranges, route tables, and gateways similar to an on-prem network. Use a VPC to isolate resources for security, place databases in private subnets, and connect securely to your on-prem network via VPN/Direct Connect.
Q: Why is IAM important in AWS?
A: IAM (Identity and Access Management) controls access to your AWS environment. It lets you create users, groups, and roles with precise permissions. By using IAM, you avoid using the root account and practice least privilege, minimizing risk. For example, give an EC2 instance a role with only the S3 permissions it needs rather than embedding keys on the server.
Q: When should I use Amazon EC2 vs AWS Lambda?
A: Use EC2 when you need full control of a server (long-running processes, complex applications, any OS-level software). Use Lambda for event-driven, short-lived tasks (microservices, file processing, scheduled jobs) where you don’t want to manage servers. Lambda scales automatically per request, while EC2 gives you a full VM environment.
Q: How do Amazon S3 and EBS differ?
A: S3 is object storage: you store files ("objects") in buckets. It’s web-accessible and virtually unlimited. It’s best for static assets, backups, or data lakes. EBS (Elastic Block Store) is like a traditional disk attached to an EC2 instance - it’s block storage. EBS is for the filesystem needs of a running server (e.g., your OS and app on EC2 would boot from an EBS volume).
Q: What is CloudFront and why use it?
A: CloudFront is a global CDN (Content Delivery Network) (aws.amazon.com). It caches your content at edge locations around the world. Users fetch data (like images or API responses) from an edge node near them instead of your origin server, reducing latency. It also reduces load on your origin. In short, use CloudFront to speed up delivery of content to worldwide users.
Q: How does Amazon Route 53 differ from other DNS services?
A: Route 53 is AWS’s DNS via the cloud. It supports advanced routing policies (like weighted, latency-based, failover) beyond basic name-to-IP (which most DNS providers do). It also integrates with AWS health checks: Route 53 can automatically remove a server from DNS if it fails. Additionally, you can register and manage domain names directly in Route 53.
Q: What is CloudWatch, and do I need CloudWatch Logs?
A: CloudWatch handles monitoring and logging on AWS. It collects metrics (CPU, memory, etc.) and can trigger alarms. CloudWatch Logs is used for collecting and searching log data (from EC2, Lambda, etc.). Yes, you typically use CloudWatch Logs to aggregate logs in one place and set metrics/alarms on them. It’s an essential tool for diagnosing issues and ensuring your system is running healthily.
Q: How do SNS and SQS differ?
A: SNS (Simple Notification Service) is push-based pub/sub. You publish a message to a topic and SNS pushes it immediately to all subscribers (email, Lambda, HTTP endpoints, SQS queues, etc.). SQS (Simple Queue Service) is pull-based messaging: producers post messages to a queue, and consumers poll the queue to retrieve messages. Use SNS for instant fan-out notifications; use SQS for decoupling components via a reliable queue.
Q: How should I get started with AWS after learning these basics?
A: Practice by building simple projects: launch an EC2-backed web server serving a site from S3 with a CloudFront CDN and Route 53 DNS. Add CloudWatch alarms. Secure your setup with proper IAM roles. Then explore more specialized services (containers, machine learning, etc.). Consider official AWS certification study or hands-on training like the Refonte Cloud Engineer Program to solidify your knowledge and skills.
