Change Application Capacity with a Scaling Policy

AWSBeginner
Practice Now

Introduction

An application sometimes needs more servers and later needs fewer. In this lab, you will define two scaling policies, expand a real fleet from two to three instances and return it to two. You will confirm actual HTTP responses and see how group bounds prevent further expansion or reduction.

You should understand launch templates, Auto Scaling groups and ALB health checks. This fresh environment provides an independent, healthy two-instance group named application-fleet, its template, network and load balancer. No resources from earlier labs are reused.

Certification Relevance

This lab supports introductory practice for SAA-C03 Domain 2 and SOA-C03 Domain 2: configuring Auto Scaling capacity, scaling actions and load balancer integration.

Define the Scaling Policies

In this step, you will inspect the supplied fleet and create policies without changing capacity.

Enter the workspace and inspect the group's starting capacity:

cd /home/labex/project

aws autoscaling \
  describe-auto-scaling-groups \
  --auto-scaling-group-names application-fleet \
  --query 'AutoScalingGroups[].{Min:MinSize,Desired:DesiredCapacity,Max:MaxSize,Instances:Instances[].InstanceId}'

TG_ARN=$(aws elbv2 \
  describe-target-groups \
  --names application-targets \
  --query 'TargetGroups[0].TargetGroupArn' \
  --output text)

LB_DNS=$(aws elbv2 \
  describe-load-balancers \
  --names application-alb \
  --query 'LoadBalancers[0].DNSName' \
  --output text)

aws elbv2 \
  wait target-in-service \
  --target-group-arn "$TG_ARN"

Expect minimum 2, desired 2, maximum 3 and two healthy targets in AWS View. A scaling policy defines a capacity adjustment. ChangeInCapacity adds or removes an absolute number: 1 adds one server; -1 removes one.

Create two SimpleScaling policies. A cooldown normally allows a simple scaling action to settle before another alarm-driven action. Store 30 seconds here:

aws autoscaling \
  put-scaling-policy \
  --auto-scaling-group-name application-fleet \
  --policy-name add-one-server \
  --policy-type SimpleScaling \
  --adjustment-type ChangeInCapacity \
  --scaling-adjustment 1 \
  --cooldown 30

aws autoscaling \
  put-scaling-policy \
  --auto-scaling-group-name application-fleet \
  --policy-name remove-one-server \
  --policy-type SimpleScaling \
  --adjustment-type ChangeInCapacity \
  --scaling-adjustment -1 \
  --cooldown 30

aws autoscaling \
  describe-policies \
  --auto-scaling-group-name application-fleet \
  --query 'ScalingPolicies[].{Name:PolicyName,Type:PolicyType,Adjustment:ScalingAdjustment,Cooldown:Cooldown}'

Creating policies does not execute them. This lab invokes them explicitly with --no-honor-cooldown so you can inspect controlled transitions. It does not demonstrate elapsed cooldown enforcement or a CloudWatch alarm. In production, metric-triggered policies use alarms; target tracking follows a target metric value. Neither automatic metric feedback nor CPU load is tested here.

Scale Out to Three Servers

In this step, you will execute the positive policy and verify that the additional server actually serves requests.

Scale-out adds instances. Execute the policy and wait for ALB readiness:

aws autoscaling \
  execute-policy \
  --auto-scaling-group-name application-fleet \
  --policy-name add-one-server \
  --no-honor-cooldown

aws elbv2 \
  wait target-in-service \
  --target-group-arn "$TG_ARN"

aws autoscaling \
  describe-auto-scaling-groups \
  --auto-scaling-group-names application-fleet \
  --query 'AutoScalingGroups[].{Desired:DesiredCapacity,Max:MaxSize,Instances:Instances[].InstanceId}'

Expect desired capacity 3 and three current IDs. The group uses its launch template for the new server and automatically registers it with the target group.

Execute the same policy once more at the maximum, then inspect capacity again:

aws autoscaling \
  execute-policy \
  --auto-scaling-group-name application-fleet \
  --policy-name add-one-server \
  --no-honor-cooldown

aws autoscaling \
  describe-auto-scaling-groups \
  --auto-scaling-group-names application-fleet \
  --query 'AutoScalingGroups[].{Desired:DesiredCapacity,Max:MaxSize,Instances:Instances[].InstanceId}'

Capacity stays at 3; a scaling adjustment cannot exceed the group's maximum. This bound limits fleet size, not the number of requests each server can process.

Send nine requests and compare the returned IDs with the three current instances:

for request in 1 2 3 4 5 6 7 8 9; do
  curl --config client.conf -sS "http://$LB_DNS/health"
  echo
done

Open AWS View. Confirm desired 3, three healthy targets and actual 200 responses from all three IDs using Send request. Adding a resource in an API response alone would not prove it serves traffic.

Three serving instances

Example: desired capacity is three, and the additional server returns HTTP 200. Your instance IDs will differ.

Scale In and Preserve the Minimum

In this step, you will reduce capacity and confirm that two application servers remain available.

Scale-in removes instances. Run the negative policy and wait for current targets to be healthy:

aws autoscaling \
  execute-policy \
  --auto-scaling-group-name application-fleet \
  --policy-name remove-one-server \
  --no-honor-cooldown

aws elbv2 \
  wait target-in-service \
  --target-group-arn "$TG_ARN"

aws autoscaling \
  describe-auto-scaling-groups \
  --auto-scaling-group-names application-fleet \
  --query 'AutoScalingGroups[].{Min:MinSize,Desired:DesiredCapacity,Instances:Instances[].InstanceId}'

aws ec2 \
  describe-instances \
  --filters "Name=tag:aws:autoscaling:groupName,Values=application-fleet" \
  --query 'Reservations[].Instances[].{Instance:InstanceId,State:State.Name}'

Expect two current IDs and one terminated instance. The group chooses which instance to remove; do not assume the newest server is always selected. The removed instance no longer belongs in the target group's serving set.

Execute the negative policy again at the minimum, then inspect capacity:

aws autoscaling \
  execute-policy \
  --auto-scaling-group-name application-fleet \
  --policy-name remove-one-server \
  --no-honor-cooldown

aws autoscaling \
  describe-auto-scaling-groups \
  --auto-scaling-group-names application-fleet \
  --query 'AutoScalingGroups[].{Min:MinSize,Desired:DesiredCapacity,Instances:Instances[].InstanceId}'

Desired capacity stays at 2; the policy cannot reduce the group below its minimum. Send six requests to confirm that the remaining fleet works:

for request in 1 2 3 4 5 6; do
  curl --config client.conf -sS "http://$LB_DNS/health"
  echo
done

In AWS View, confirm two healthy targets and 200 responses from both current IDs. The terminated ID must not appear in new responses. This exercise checks lifecycle and routing; it does not benchmark scale-in draining time or application session preservation.

Remove the Policies and Keep the Prepared Fleet

In this step, you will delete your policy configuration after returning the supplied fleet to two servers.

Delete both policies and inspect the remaining policy inventory:

aws autoscaling \
  delete-policy \
  --auto-scaling-group-name application-fleet \
  --policy-name add-one-server

aws autoscaling \
  delete-policy \
  --auto-scaling-group-name application-fleet \
  --policy-name remove-one-server

aws autoscaling \
  describe-policies \
  --auto-scaling-group-name application-fleet \
  --query 'ScalingPolicies[].PolicyName'

aws autoscaling \
  describe-auto-scaling-groups \
  --auto-scaling-group-names application-fleet \
  --query 'AutoScalingGroups[].{Min:MinSize,Desired:DesiredCapacity,Max:MaxSize}'

The policy list should be [], and group capacity should remain minimum 2, desired 2, maximum 3. Policy deletion does not terminate the existing fleet. Keep the supplied group, template, load balancer and network in place; AWS View should still show two healthy servers.

Summary

You defined positive and negative simple scaling policies, observed real scale-out from two to three servers and scale-in back to two, and tested both capacity bounds. You checked actual HTTP responses and native instance termination, then removed the policies while preserving the restored fleet.

You can now distinguish policy configuration, explicit execution and metric-driven triggering. The challenge applies the course's routing and health-check diagnosis skills to a new server that does not receive traffic.