Cold Start
- In Turkish
- soğuk başlangıç
In short
A cold start is the extra delay when a serverless platform must create a new function instance, start its runtime and run setup code before handling a request.
What is a cold start?
Serverless platforms such as AWS Lambda, Azure Functions and Google Cloud Run don't keep your code running all the time. When a request arrives and no instance is ready, the platform has to create one: it allocates a lightweight virtual machine or container, downloads the code or image, starts the runtime, such as Node.js, Python or the Java virtual machine, and runs your initialization code. Only then can it handle the request. That extra wait is the cold start.
Once an instance exists, the platform keeps it for a while and reuses it, so the next requests are warm starts that skip the setup. Cold starts happen whenever a new instance is needed: after the function has been idle, after a new deployment, and when traffic grows and the platform scales out. AWS says they typically affect under 1% of Lambda invocations and last from under 100 milliseconds to over a second; large packages, heavy frameworks and runtimes such as Java and .NET tend to take longest.
Teams shorten cold starts by keeping deployment packages small, loading libraries only when they are needed and creating clients such as database connections once, outside the request handler, so warm requests reuse them. Platforms also offer paid ways to keep instances ready: provisioned concurrency on AWS Lambda, minimum instances on Cloud Run and always-ready instances on Azure Functions. Lambda SnapStart, available for Java, Python and .NET, starts new instances from a snapshot of an already initialized one. Platforms built on V8 isolates, such as Cloudflare Workers, start code in a few milliseconds, so their cold starts are barely noticeable.
A cold start is often confused with general slowness. It only adds time to the first request handled by a new instance; if every request is slow, the cause lies elsewhere, in the code, the database or the network. The term is also used outside serverless, for an app launching from scratch or a cache that starts empty, but the idea is the same: the first use pays for setup that later uses skip.
Key takeaways
- A cold start is the setup delay when a platform must start a new function instance.
- Warm starts reuse an existing instance and skip that setup.
- Cold starts follow idle periods, new deployments and scaling out.
- Small packages and initialization outside the handler make them shorter.
- Provisioned concurrency and minimum instances keep instances ready; SnapStart makes new ones start faster.
Example
import boto3
# Runs once per new instance, during the cold start
s3 = boto3.client("s3")
def handler(event, context):
# Runs on every request; warm starts reuse the client created above
obj = s3.get_object(Bucket="reports", Key=event["key"])
return {"size": obj["ContentLength"]}Readers ask
How long does a cold start take?
It depends on the platform, the runtime and the size of the code. AWS describes Lambda cold starts as lasting from under 100 milliseconds to over a second; small Node.js or Python functions sit at the fast end, and large Java or .NET applications without SnapStart at the slow end.
How can I avoid cold starts?
You can make them shorter by keeping packages small and initialization light, and avoid most of them by paying for instances that are always ready, such as provisioned concurrency on AWS Lambda or minimum instances on Cloud Run. Sending regular warm-up requests is an older workaround that doesn't help when traffic needs more instances.
Does every request have a cold start?
No. Only a request that has to be handled by a new instance does. Requests that arrive while an instance is still alive are warm starts and skip the setup.
See also
- ServerlessDevOps & Cloud, p. 54Serverless is a cloud model in which the provider runs your code on demand, manages all the servers, scales automatically, and bills only for actual use.
- LatencyNetworking, p. 15Latency is the delay between sending a request and the start of a response, usually measured in milliseconds, and it shapes how responsive an app feels.
- AutoscalingDevOps & Cloud, p. 2Autoscaling is the automatic adding or removing of computing resources, such as servers or containers, based on demand to keep performance steady and costs low.
- Edge ComputingDevOps & Cloud, p. 24Edge computing runs code and processes data close to where users or devices are, instead of in a distant central data center, to reduce latency.
- JVMProgramming Languages, p. 17The JVM runs Java bytecode, so the same compiled program works on any operating system; Kotlin, Scala and Clojure compile to that bytecode too.
- PaaSDevOps & Cloud, p. 46PaaS (platform as a service) is a cloud model where you deploy your code and the provider runs everything under it: servers, operating systems and scaling.
Sources
- AWS Lambda documentation: Understanding the Lambda execution environment lifecycledocs.aws.amazon.com(opens in a new tab)
- AWS Lambda documentation: Configuring provisioned concurrency for a functiondocs.aws.amazon.com(opens in a new tab)
- Cloud Run documentation: Set minimum instances for servicesdocs.cloud.google.com(opens in a new tab)
Spotted a mistake or something missing on this page?Suggest an edit