Articles

A Hands-On Guide to AWS Spot Instances

AWS Spot Instances sell unused EC2 capacity at a steep discount. Here's how to launch one, handle interruptions, and pick workloads that fit.

Chisato Chisato · · 4 min read
Server racks with cables in a data center

AWS Spot Instances let you run EC2 compute at a steep discount compared to on-demand pricing by using spare capacity AWS isn’t currently selling at full price — in exchange, AWS can reclaim that capacity with as little as two minutes’ notice if it’s needed elsewhere. This walkthrough covers how to launch one, how to handle that interruption gracefully, and which workloads are actually a good fit.

Why spot capacity is cheap

Every AWS availability zone has EC2 capacity that isn’t currently rented out on-demand or reserved. Rather than let it sit idle, AWS auctions it off as spot capacity at a price that floats with supply and demand for that instance type in that zone. You’re not bidding against other users directly the way a Dutch auction or spot pricing used to work years ago — today you set a maximum price you’re willing to pay (defaulting to the on-demand price if you don’t), and as long as the current spot price stays under that ceiling, your instance keeps running.

The catch is the trade AWS is making: if on-demand demand for that capacity picks up, or the spot price for your instance type rises above what you’re willing to pay, your instance can be interrupted. AWS gives a two-minute warning via an instance metadata endpoint and, when available, an EventBridge notification before terminating or stopping it.

Step 1: pick a workload that tolerates interruption

Before touching the console, decide whether spot actually fits what you’re running. Good fits: batch data processing jobs that checkpoint progress, CI/CD build agents, rendering or transcoding pipelines, stateless web servers behind a load balancer with enough replicas that losing one doesn’t matter, and machine learning training jobs that checkpoint periodically. Poor fits: anything stateful without its own replication (a single database instance), anything with a hard latency SLA that can’t tolerate a mid-request interruption, and anything where losing hours of uncheckpointed work would be expensive.

Step 2: launch a spot instance from the console or CLI

From the EC2 console, when launching an instance, the purchasing option section lets you select “Spot Instances” instead of the on-demand default. You’ll be asked to set a maximum price (or accept the on-demand price as the ceiling, which is the simplest default and still gets you the discount, since you only ever pay the current spot price, not your ceiling) and an interruption behavior — stop, hibernate, or terminate.

The equivalent from the AWS CLI:

aws ec2 run-instances \
  --image-id ami-0abcdef1234567890 \
  --instance-type m6i.large \
  --instance-market-options '{
    "MarketType": "spot",
    "SpotOptions": {
      "MaxPrice": "0.05",
      "SpotInstanceType": "one-time",
      "InstanceInterruptionBehavior": "terminate"
    }
  }' \
  --key-name my-key \
  --subnet-id subnet-0123456789abcdef0

SpotInstanceType set to one-time requests a single instance; persistent keeps re-requesting capacity if your instance is interrupted, which is useful for long-running workloads that just need a spot instance running, not necessarily the same one.

Step 3: handle the two-minute interruption notice

This is the part that determines whether spot actually saves you money or just causes outages. Every EC2 instance can poll its own instance metadata service for an interruption notice:

curl -s http://169.254.169.254/latest/meta-data/spot/instance-action

This returns nothing under normal operation, and returns a JSON object with an action (stop, hibernate, or terminate) and a time once AWS has scheduled an interruption — which, by design, comes roughly two minutes ahead of it happening. A minimal handler polls this endpoint every five to ten seconds and, on seeing a response, does whatever cleanup is appropriate: flush a checkpoint, deregister from a load balancer’s target group, drain in-flight requests, or push a final log batch. For workloads managed by an orchestrator, this handling is often already built in — an Auto Scaling group with mixed instance types, or a Kubernetes cluster running spot node groups (typically marked with taints so only tolerant pods land on them), will drain and reschedule work automatically on the interruption signal rather than requiring you to write the polling loop yourself.

Step 4: build in redundancy, not just discount

Spot pricing and availability vary independently across instance types and availability zones, so the most reliable pattern is requesting a range of acceptable instance types rather than one specific one — an m6i.large might be interrupted while an m6a.large in the same zone has capacity to spare. AWS’s own recommendation, and the default behavior of EC2 Fleet and Auto Scaling groups configured for spot, is to diversify across several instance types and zones simultaneously rather than betting on a single configuration staying available. Full details on interruption behavior, pricing history, and fleet configuration are in AWS’s Spot Instances documentation.

What savings actually look like

Spot discounts vary by instance type, region, and real-time demand rather than following a fixed percentage, but the durable pattern is that spot capacity for widely available, less popular instance types tends to be both cheaper and less frequently interrupted than spot capacity for the newest, most in-demand instance families — everyone else is also trying to grab the popular ones on-demand.

The takeaway

Spot Instances trade a real risk of interruption for a substantial discount on EC2 compute, which is a good trade specifically for stateless, checkpointable, or horizontally redundant workloads — and a bad one for anything that can’t tolerate losing an instance with two minutes’ notice. The mechanics that make it safe in practice are the interruption-notice polling (or letting an orchestrator handle it) and diversifying across instance types and zones rather than depending on a single configuration staying available.

Chisato Chisato · · 4 min read

Incident Severity Levels Explained (SEV1-SEV4)

Incident severity levels rank outages by impact so teams respond proportionally. What SEV1 through SEV4 typically mean and how to set the scale.

#DevOps #Cloud #Observability
Chisato Chisato · · 4 min read

Kubernetes Init Containers Explained

Init containers run to completion before a pod's main containers start, making them the standard way to handle setup steps and startup ordering.

#Kubernetes #DevOps #Cloud