Maintain Capacity with an Auto Scaling Group

AWSBeginner
Practice Now

Introduction

A load balancer can avoid a failed application, but it does not by itself create a replacement server. In this lab, you will create a launch template and an Auto Scaling group that maintains two application instances. You will cause one application to fail, observe automatic replacement with a new instance ID and verify that the new server actually serves requests.

You should understand EC2 User Data, ALB target groups and application health checks. This fresh environment supplies the network, application image, key pair and empty ALB target group with an HTTP listener. It does not contain an existing application fleet. The AWS CLI and connection files are prepared independently of earlier labs.

Certification Relevance

This lab provides introductory hands-on practice for the following exam topics.

Create the Application Launch Template

In this step, you will define the configuration that Auto Scaling uses to launch each application server.

Enter the workspace and load the prepared network variables:

cd /home/labex/project
source launch.env

A launch template stores instance settings such as AMI, type, security groups and startup configuration. Inspect the supplied payload:

cat launch-template.json
cat application-user-data.sh

The JSON selects the prepared AMI, t3.micro, report-key and the application security group. Its UserData contains the base64-encoded startup script shown separately. The script creates the application's initial configuration with message Application ready and healthy set to true. Auto Scaling must apply this at every launch so replacements are ready to serve; a manual change inside one old server does not change the template.

Create your template. file:// tells AWS CLI to read the JSON file as the parameter value:

LT_ID=$(aws ec2 \
  create-launch-template \
  --launch-template-name application-template \
  --launch-template-data file://launch-template.json \
  --query 'LaunchTemplate.LaunchTemplateId' \
  --output text)

Templates have numbered versions. Inspect version 1, which your group will use explicitly:

aws ec2 \
  describe-launch-template-versions \
  --launch-template-id "$LT_ID" \
  --versions 1 \
  --query 'LaunchTemplateVersions[].{Version:VersionNumber,AMI:LaunchTemplateData.ImageId,Type:LaunchTemplateData.InstanceType,Key:LaunchTemplateData.KeyName,Groups:LaunchTemplateData.SecurityGroupIds}'

Creating a template does not launch instances. Open AWS View: the supplied ALB and target group exist, but there are no registered application targets yet.

Launch and Connect the Application Fleet

In this step, you will create an Auto Scaling group and connect its actual application instances to the supplied ALB.

An Auto Scaling group maintains a desired number of instances within its minimum and maximum. The desired capacity is the count it works to maintain, not a measurement of request throughput. Here, minimum 2, desired 2 and maximum 3 provide a two-server starting fleet with room for one additional server.

Capture the supplied target group ARN and load balancer DNS name:

TG_ARN=$(aws elbv2 \
  describe-target-groups \
  --names application-targets \
  --query 'TargetGroups[0].TargetGroupArn' \
  --output text)

LB_DNS=$(aws elbv2 \
  describe-load-balancers \
  --names application-alb \
  --query 'LoadBalancers[0].DNSName' \
  --output text)

Create the group using template version 1 and both prepared subnets. --target-group-arns connects group instances to ALB registration automatically. --health-check-type ELB includes load balancer health in replacement decisions; EC2 checks alone do not identify every application failure. The grace period gives new instances time to start before application health can cause replacement. This exercise uses 30 seconds; production values must match startup behavior:

aws autoscaling \
  create-auto-scaling-group \
  --auto-scaling-group-name application-fleet \
  --launch-template "LaunchTemplateId=$LT_ID,Version=1" \
  --min-size 2 \
  --max-size 3 \
  --desired-capacity 2 \
  --vpc-zone-identifier "$SUBNET_ID,$SECOND_SUBNET_ID" \
  --target-group-arns "$TG_ARN" \
  --health-check-type ELB \
  --health-check-grace-period 30

Inspect the group and its instance IDs:

aws autoscaling \
  describe-auto-scaling-groups \
  --auto-scaling-group-names application-fleet \
  --query 'AutoScalingGroups[].{Min:MinSize,Desired:DesiredCapacity,Max:MaxSize,HealthCheck:HealthCheckType,Grace:HealthCheckGracePeriod,Instances:Instances[].InstanceId}'

Wait for application readiness, then inspect the ALB's targets:

aws elbv2 \
  wait target-in-service \
  --target-group-arn "$TG_ARN"

aws elbv2 \
  describe-target-health \
  --target-group-arn "$TG_ARN" \
  --query 'TargetHealthDescriptions[].{Instance:Target.Id,Health:TargetHealth.State}' \
  --output table

Expect two healthy IDs. Send six requests and compare the response IDs with the group's instances:

for request in 1 2 3 4 5 6; do
  curl --config client.conf -sS "http://$LB_DNS/health"
  echo
done

Open AWS View. Confirm desired capacity 2, two healthy targets and actual 200 responses from both IDs using Send request. Subnet and Availability Zone labels show configuration; this exercise does not measure physical cross-zone resilience.

Observe Automatic Failure Replacement

In this step, you will cause an application failure and let the group replace that instance instead of repairing it manually.

Select one current group instance and capture its public address for SSH:

OLD_ID=$(aws autoscaling \
  describe-auto-scaling-groups \
  --auto-scaling-group-names application-fleet \
  --query 'AutoScalingGroups[0].Instances[0].InstanceId' \
  --output text)

OLD_IP=$(aws ec2 \
  describe-instances \
  --instance-ids "$OLD_ID" \
  --query 'Reservations[0].Instances[0].PublicIpAddress' \
  --output text)

Make its application return 503 by changing only healthy. As in the preceding lab, a temporary JSON file avoids overwriting the file during reading:

ssh -F ssh_config "ubuntu@$OLD_IP" \
  'sudo jq ".healthy = false" /etc/report-app/config.json > /tmp/report-config.json && sudo install -m 644 /tmp/report-config.json /etc/report-app/config.json'

ssh -F ssh_config "ubuntu@$OLD_IP" \
  'curl -s -o /dev/null -w "%{http_code}\n" http://127.0.0.1:8081/health'

Expect 503. The application failed while the instance was running. The ALB detects failed health checks; after the startup grace period, the group replaces the unhealthy instance to restore its desired capacity. The same launch template starts the replacement with a healthy application configuration.

Keep AWS View open to observe the changing IDs and health states. An unhealthy state can be brief because replacement follows detection; a new target can appear as initial before becoming healthy. Do not issue a repair or an extra run-instances command.

Wait for the old instance to be terminated and for the target group to be healthy:

aws ec2 \
  wait instance-terminated \
  --instance-ids "$OLD_ID"

aws elbv2 \
  wait target-in-service \
  --target-group-arn "$TG_ARN"

Inspect the old instance and the current group:

aws ec2 \
  describe-instances \
  --instance-ids "$OLD_ID" \
  --query 'Reservations[].Instances[].{Instance:InstanceId,State:State.Name}'

aws autoscaling \
  describe-auto-scaling-groups \
  --auto-scaling-group-names application-fleet \
  --query 'AutoScalingGroups[].{Desired:DesiredCapacity,Instances:Instances[].InstanceId}'

The old ID should be terminated. Desired capacity remains 2, with a new ID replacing the old one. A replacement is a newly initialized server; changes made only inside the old server are not preserved automatically.

Send six requests again:

for request in 1 2 3 4 5 6; do
  curl --config client.conf -sS "http://$LB_DNS/health"
  echo
done

Look for both current group IDs, including the replacement, and no old ID. In AWS View, confirm two healthy targets and use Send request to observe the replacement actually responding. Native resource counts alone would not prove that the new application works.

Replacement instance serving

Example: desired capacity stays at two, and the newly created instance returns HTTP 200. Your instance IDs will differ.

Delete the Fleet and Launch Template

In this step, you will remove your group, its instances and its launch template, while preserving the supplied ALB and network.

Delete the group with --force-delete to terminate its instances even though its minimum capacity is 2. Deleting a template alone would not stop running instances:

aws autoscaling \
  delete-auto-scaling-group \
  --auto-scaling-group-name application-fleet \
  --force-delete

aws ec2 \
  wait instance-terminated \
  --filters "Name=tag:aws:autoscaling:groupName,Values=application-fleet"

aws ec2 \
  delete-launch-template \
  --launch-template-id "$LT_ID"

The waiter selects this group's instances by the tag that Auto Scaling adds automatically. It includes the old replaced instance and waits for all selected instances to terminate. Check the group's absence, the template inventory and target registrations:

aws autoscaling \
  describe-auto-scaling-groups \
  --auto-scaling-group-names application-fleet \
  --query 'AutoScalingGroups[].AutoScalingGroupName'

aws ec2 \
  describe-launch-templates \
  --query 'LaunchTemplates[].LaunchTemplateName'

aws elbv2 \
  describe-target-health \
  --target-group-arn "$TG_ARN" \
  --query 'TargetHealthDescriptions[].Target.Id'

All three lists should be []. Target deregistration can take time; if it is still pending, wait briefly and repeat the last query. AWS View should retain the supplied ALB and target group but show no fleet and no registered application targets. Leave the supplied network, AMI and key pair in place.

Summary

You created a versioned launch template and an Auto Scaling group that maintained two real application servers. You included ALB health checks in replacement decisions, caused an application failure and verified that a new instance ID actually served requests. Finally, you removed the group, its instances and the template while preserving the supplied load balancer and network.

The next lab changes desired capacity through simple scaling policies and tests the group bounds.