Article
Using AWS Lambda scalable bandwidth: weigh transfer latency against memory cost
Eligible Lambda functions outside your VPC can scale bandwidth with memory after quota approval. Parallel S3 reads and an AWS benchmark show why lower latency and lower execution cost must be evaluated separately.
Share
Koharu's reading tip
Read latency and cost as separate questions. The key distinction is between the time a user waits and the sum of billed time across workers.

A Lambda function that reads large datasets can spend much of its time waiting for transfers, even when computation is fast. The AWS bandwidth explainer published on September 28, 2026 offers a reason to revisit that wait through memory configuration.
Does allocating more memory necessarily make the function faster and cheaper? Connecting eligibility, S3 access patterns, and duration pricing helps identify which functions to test and which numbers to compare.
Lambda bandwidth scaling starts with quota approval
Scalable bandwidth applies to functions not attached to your VPC and configured with at least 2,048 MB of memory. Sustained bandwidth per execution environment can reach 3,000 Mbps at 10,240 MB. Changing memory alone does not enable it.
First, request an increase for Network bandwidth per execution environment in Service Quotas. Memory-based scaling becomes available after approval. The Lambda quotas documentation describes these conditions.
The feature was announced on August 5, 2026 with no additional charge in commercial AWS Regions; the September article develops its practical use and measurements. No additional feature charge does not mean that increasing memory leaves execution charges unchanged.
Outside a VPC here means not attached to your own VPC. Lambda normally runs functions within an AWS-managed VPC. As the networking documentation explains, attachment provides access to private resources. Bandwidth alone is therefore insufficient grounds to remove a connection that the application needs.
Parallel S3 reads put the extra bandwidth to work
A higher ceiling helps only if the application can supply enough transfer work. The S3 performance guidelines recommend spreading requests across connections and fetching different byte ranges of an object concurrently.
The AWS sample implementation supports both pre-partitioned objects distributed across workers and byte ranges of one large object. Distinguish parallelism across function instances from parallel transfers inside each instance.
For Python, managed Boto3 transfers such as download_fileobj can use the AWS Common Runtime (CRT). The Boto3 transfer reference documents preferred_transfer_client="crt" in TransferConfig. This setting does not turn low-level get_object calls into managed parallel transfers. The byte-range sample controls concurrent requests explicitly.
When testing higher concurrency, watch S3 503 responses and retries alongside throughput. More connections should reduce waiting; rising failures are a reason to reassess whether the setting improves the complete operation.
A 10 GB benchmark separates bandwidth from query latency
The AWS measurement uses 20 work items of 512 MB each, totaling 10 GB. Changing worker memory produced these end-to-end wall times. The published benchmark is an example, not a performance guarantee for another environment.
| Worker memory | Query wall time, p50 |
|---|---|
| 1,024 MB | 7.113 seconds |
| 10,240 MB | 2.640 seconds |
Median latency fell by approximately 62.9%. The maximum bandwidth ratio, 3,000 Mbps versus the previous sustained ceiling of 625 Mbps, is 4.8, but application speed need not increase by that same factor.
CPU allocation changes too. Lambda memory configuration increases CPU proportionally, so this comparison does not isolate network improvements. Recording data retrieval separately from transformation and aggregation can help identify the next bottleneck in your own workload.
Choose memory by comparing latency with GB-seconds
Lambda duration pricing uses allocated memory and billed duration, expressed in GB-seconds. Holding Region, architecture, and unit price constant, the basic relationship for duration charges is:
Duration charge ratio ≈ Memory ratio × Billed duration ratio
For example, raising memory tenfold from 1,024 MB to 10,240 MB requires total billed duration to fall below one tenth, at the same invocation count, to reduce duration charges. This is arithmetic from the pricing model, not a cost estimate for the benchmark above. Parallel query wall time is different from the sum of worker billed durations.
Start by recording a baseline at the current memory setting, then measure after enablement with the same data and concurrency. Next, vary memory in steps and compare user-observed latency, billed duration per invocation, and total GB-seconds. Separate first invocations from reused environments and inspect the slow end of the distribution as well as p50.
Functions dominated by transfer waits have a reason to test scalable bandwidth. The best setting need not be maximum memory. Compare the total cost of memory and parallelism configurations that meet the required response time: that is how this change becomes an operational decision.
Source
- Title: Improving Lambda function latency with scalable network bandwidth
- URL: https://aws.amazon.com/blogs/compute/improving-lambda-function-latency-with-scalable-network-bandwidth/
Share
Related Articles
These articles share nearby categories or tags, so you can keep reading along the same thread.




