VPC endpoints vs NAT gateways on AWS
Instances in a private subnet have no route to an internet gateway, yet they almost always need to reach something outside the VPC: Amazon S3, DynamoDB, an AWS API such as Secrets Manager, a partner's service, or an operating system package mirror on the internet. AWS gives you three ways out: a NAT gateway, a gateway endpoint and an interface endpoint. They differ in what they can reach, what they cost and how tightly you can lock them down.
The exams turn this into cost and compliance puzzles. A question describes a big NAT bill, or a rule that says traffic to S3 must never touch a NAT device, or a SaaS provider that needs to expose one service to hundreds of customer VPCs with overlapping CIDR ranges. All three options can carry the traffic, so the constraint in the question decides which one is right.
NAT gateway
A NAT gateway lets instances in private subnets start connections to destinations outside the VPC while blocking connections started from outside. A public NAT gateway sits in a public subnet with an Elastic IP address and sends traffic through the internet gateway. A private NAT gateway has no Elastic IP and is used to reach other VPCs or on-premises networks through a transit gateway or virtual private gateway.
Facts worth knowing:
- It is billed per hour and per gigabyte processed, on top of any normal data transfer charges. Large flows to S3 through a NAT gateway pay processing on every byte.
- It starts at 5 Gbps and scales automatically to 100 Gbps.
- You can't attach a security group to it. You control traffic with the instances' security groups and the subnet's network ACLs.
- A standard (zonal) NAT gateway lives in one Availability Zone and is redundant only inside that zone. If private subnets in three zones all route to one gateway, losing that zone cuts outbound access for all three. The fix is one gateway per zone, each private subnet routing to the gateway in its own zone. That also avoids paying cross-zone data transfer.
- AWS now also offers regional NAT gateways, which expand across zones automatically and don't need a public subnet. They don't support private NAT. Exam questions still mostly assume the zonal model, so read the options carefully.
A NAT gateway is the only one of the three that reaches arbitrary internet destinations. Package updates, third-party APIs and anything else that isn't an AWS or PrivateLink service still need it (or an egress-only internet gateway for IPv6).
Gateway endpoints
A gateway endpoint exists for exactly two services: Amazon S3 and DynamoDB. It doesn't use PrivateLink and has no additional charge.
You don't connect to it by IP address. When you associate a gateway endpoint with route tables, AWS adds a route whose destination is the service's AWS-managed prefix list and whose target is the endpoint. Routing picks the most specific match, so for that service in the same Region this route beats a 0.0.0.0/0 route to a NAT gateway or internet gateway. Everything else keeps using the default route.
Its limits are what the exams test:
- It works only from inside the VPC. Traffic arriving over Direct Connect, Site-to-Site VPN, VPC peering or a transit gateway can't use it.
- It reaches only the service in the same Region. A bucket in another Region still goes out the default route.
- Only subnets whose route tables are associated with it use it. Other subnets keep using the public endpoint.
- Security groups and network ACLs must still allow the traffic. Security groups can reference the prefix list. Network ACLs can't, so you add the service's CIDR ranges instead.
Interface endpoints and PrivateLink
An interface endpoint places an elastic network interface with a private IP address in each subnet you choose, one subnet per Availability Zone. Traffic to the service goes to those IPs and travels over AWS PrivateLink. Most AWS services support interface endpoints, and S3 and DynamoDB support both types.
- It is billed per hour per Availability Zone and per gigabyte processed.
- It has a security group, so you can control which sources may use it.
- With private DNS enabled, the service's normal public hostname resolves to the endpoint's private IPs inside the VPC, so SDKs and applications work unchanged.
- Because it is just IP addresses in your VPC, it can be reached from on premises over Direct Connect or VPN and from peered or transit-gateway-connected VPCs. On-premises DNS needs a Route 53 Resolver inbound endpoint or the endpoint-specific DNS names.
- For high availability, put it in at least two zones.
PrivateLink also works for your own services. A provider puts an application behind a Network Load Balancer and publishes an endpoint service. Consumers create an interface endpoint in their own VPC. Traffic only flows from the consumer to that one service, so the consumer can reach nothing else in the provider's VPC, and the two VPCs can have overlapping CIDR ranges because no routes are exchanged. That is the deciding difference from VPC peering.
Endpoint policies
Both endpoint types accept an endpoint policy, a resource-based policy that limits which principals, actions and resources can be used through the endpoint. A common pairing is an endpoint policy that allows only your own buckets, which stops data being copied to someone else's account. Add to that a bucket policy that denies any request whose aws:SourceVpce isn't your endpoint. Be careful with the bucket side: that deny also blocks console access, which doesn't come through the endpoint.
How to choose
- Private instances need S3 or DynamoDB in the same Region, and cost matters: gateway endpoint.
- Private instances need another AWS API (Secrets Manager, SSM, ECR, CloudWatch, STS) without internet access: interface endpoint for each service.
- On-premises servers or another Region's VPC need private access to S3: interface endpoint. A gateway endpoint can't be reached from there. Keeping a gateway endpoint as well for in-VPC traffic avoids paying for that traffic.
- You need to expose one service to many other VPCs or accounts, possibly with overlapping CIDRs: PrivateLink endpoint service behind an NLB.
- Instances need arbitrary internet destinations: NAT gateway, one per Availability Zone for resilience.
- A NAT bill is dominated by S3 or DynamoDB traffic: add gateway endpoints and keep the NAT gateway for everything else.
Common exam traps
- Picking an interface endpoint for S3 when the question says "no additional cost". It works, but it bills per hour and per gigabyte. The gateway endpoint is free.
- Expecting a gateway endpoint to serve on-premises clients. It can't be reached over VPN, Direct Connect, peering or a transit gateway.
- One NAT gateway for a multi-AZ VPC. It is a single-zone failure point, and adding Elastic IPs to it raises connection capacity, not availability.
- Using VPC peering to share one service with hundreds of customers. Peering exposes whole networks, requires non-overlapping CIDRs and doesn't scale to hundreds of peers. PrivateLink does.
- Forgetting that a NAT gateway has no security group. Options that put a security group on it describe something that doesn't exist.
- Locking a bucket to an endpoint with `aws:SourceVpce` and losing console access. The deny applies to every request that doesn't come through the endpoint, including your own console sessions.
A worked example
A company runs a data pipeline on EC2 in private subnets across three Availability Zones. Each month the instances read and write about 150 TB in S3 buckets in the same Region, pull container images from Amazon ECR, and download OS patches from public repositories. All of it goes through a single NAT gateway in one zone. The bill shows NAT data processing as a major cost, and an availability review flagged the single gateway.
The S3 traffic is the bulk of the bytes, so moving it off the NAT gateway gives most of the saving. A gateway endpoint for S3, associated with all three private route tables, does that at no extra charge. The prefix-list route wins over the default route for S3 in this Region. ECR needs interface endpoints for its API and Docker registry. Image layers are stored in S3, so they already use the gateway endpoint. The interface endpoints cost money per zone, but they keep the image pulls private. OS patches come from the internet, so a NAT gateway stays. To fix the availability finding, run one NAT gateway in each zone, with each private route table pointing at its own zone's gateway.
Replacing everything with interface endpoints would be wrong. It would put a per-gigabyte charge back on the 150 TB and still leave the patch downloads with no route.
Practice questions
- Removing NAT charges for log writes to S3
- Exposing a SaaS application to customers with overlapping CIDRs
- Outbound access lost when one Availability Zone failed
- Keeping instances to the company's own buckets through an endpoint
- NAT data processing charges across 12 VPCs