Organize Documents with Keys and Metadata

AWSBeginner
Practice Now

Introduction

A team stores finance and HR documents in one S3 bucket. You will give them meaningful names and properties, retrieve a report, and remove the finance group while preserving HR until final cleanup.

Complete Store and Retrieve Files in S3 first. This fresh VM provides the CLI connection and local documents; no earlier resources or personal credentials are needed. Use Terminal for commands and the AWS View tab beside it to compare stored keys, properties, and contents.

Certification Relevance

This lab provides hands-on practice for the following exam topics.

Choose Keys and Upload Team Documents

In this step, you will use meaningful object keys to separate two departments' documents in one bucket.

A local directory stores files on your machine; a bucket stores objects in S3. Enter the prepared workspace with cd (change directory):

cd /home/labex/project

The supplied documents directory contains two small text files. cat prints their contents so you know what you are about to store:

cat documents/revenue.csv
month,revenue
2026-09,42000

A CSV file uses commas to separate columns. Here the two columns are month and revenue; each following line is a record.

cat documents/welcome.txt
Welcome to the reporting team.

Create a bucket with aws s3 mb (make bucket). s3:// identifies a storage location rather than a local path:

aws s3 mb s3://labex-team-documents

The command reports make_bucket: labex-team-documents.

The finance key will be finance/2026-09/revenue.csv. The slashes make the name easy to group, but S3 does not create filesystem directories: the whole string is one object key. Its prefix finance/ groups finance documents; the longer prefix finance/2026-09/ narrows the group to one month.

Content-Type is standard metadata describing a file's media format. text/csv identifies comma-separated text. Custom metadata contains your own descriptive key-value pairs. Here department=finance and period=2026-09 describe the report; they do not grant permissions or change its contents.

This example separates the key’s prefix from the remaining name; metadata belongs to the object.

Keys prefix and metadata

aws s3 cp copies the first path to the second. --content-type explicitly sets the format. --metadata accepts comma-separated name=value pairs. Quoting keeps the metadata argument together. A backslash at the end of a line continues the same command on the next line:

aws s3 cp documents/revenue.csv s3://labex-team-documents/finance/2026-09/revenue.csv \
  --content-type text/csv \
  --metadata 'department=finance,period=2026-09'

The upload message names the destination key. Upload the welcome document with its own key and properties. text/plain means ordinary text without a specialized document format:

aws s3 cp documents/welcome.txt s3://labex-team-documents/hr/welcome.txt \
  --content-type text/plain \
  --metadata 'department=hr'

Use ls (list) and --recursive to display complete keys throughout the bucket instead of grouping them into folder-like prefixes:

aws s3 ls s3://labex-team-documents/ --recursive

Two lines end in finance/2026-09/revenue.csv and hr/welcome.txt; timestamps vary. AWS View shows both keys in the same bucket. You have organized objects by name, without creating separate buckets for each department.

Inspect Properties and Retrieve a Report

In this step, you will select one department's objects, inspect metadata, and prove a downloaded report matches its source.

aws s3api exposes individual S3 API operations. The higher-level aws s3 commands handle common file workflows; both operate on the same objects. list-objects-v2 lists object records. --bucket names the container and --prefix limits the keys returned by S3 to those starting with that string:

aws s3api list-objects-v2 --bucket labex-team-documents --prefix finance/

Find finance/2026-09/revenue.csv in Contents. There is no HR object in this response because hr/welcome.txt does not start with finance/. A prefix filters names, not custom metadata: department=finance alone would not place a differently named key in this result.

head-object retrieves the object's properties without downloading its body. Supply the complete key, including its prefix:

aws s3api head-object --bucket labex-team-documents --key finance/2026-09/revenue.csv

The response includes ContentType set to text/csv, ContentLength (stored bytes), and a Metadata object containing department and period. Other fields, such as timestamps and ETag, describe the stored object. Metadata is associated with the stored copy, not the original local file.

In AWS View, click finance/2026-09/revenue.csv. The expanded card shows its Content-Type, custom metadata, and actual stored CSV contents. Compare these with the CLI response.

This example shows the finance object expanded, with properties and content read from storage:

Finance key, metadata, and CSV contents

Download with the S3 URI as the source and a new local filename as the destination. Its name does not have to match the key:

aws s3 cp s3://labex-team-documents/finance/2026-09/revenue.csv retrieved-revenue.csv

Inspect what you retrieved:

cat retrieved-revenue.csv

It should show the same two CSV lines you inspected earlier. cmp compares two local files byte for byte; it produces no output on a match. The shell's && runs the message only when that comparison succeeds:

cmp documents/revenue.csv retrieved-revenue.csv && echo 'Report content matches'
Report content matches

This proves retrieval preserved the report's bytes. Your local file has the contents; use head-object when you need to inspect the properties stored in S3.

Remove Only the Finance Prefix

In this step, you will clear the finance documents while preserving the HR document.

A shared bucket can contain unrelated work. Deleting the whole bucket or using its root as a recursive deletion target would affect both departments. Choose the specific finance/ prefix instead. The trailing slash is part of the prefix and keeps names such as finance-archive.csv outside this scope.

aws s3 rm removes objects. --recursive applies the operation to all matching keys, and --dryrun lists intended operations without performing them. Preview the exact scope first:

aws s3 rm s3://labex-team-documents/finance/ --recursive --dryrun

The dry-run output names only finance/2026-09/revenue.csv. It must not name hr/welcome.txt. The objects still exist at this point.

Remove --dryrun to perform the reviewed operation:

aws s3 rm s3://labex-team-documents/finance/ --recursive

The output confirms the finance key's deletion. Query the bucket again:

aws s3 ls s3://labex-team-documents/ --recursive

Only hr/welcome.txt remains. In AWS View, expand that remaining object: its welcome message and department=hr metadata are unchanged. A successful list and read establish that you removed the intended scope while preserving another department's document.

Clean Up the Remaining Lab Resources

In this step, you will remove the remaining lab-owned document and the empty bucket.

The HR object was protected during the finance operation. Both departments' resources belong to this exercise, so you can now remove the remaining object by its exact key, without using a broad recursive target:

aws s3 rm s3://labex-team-documents/hr/welcome.txt

The output confirms the HR key's deletion. Check the bucket contents before removing its container:

aws s3 ls s3://labex-team-documents/ --recursive

The successful command prints no object rows. A bucket must be empty before rb (remove bucket) can remove it:

aws s3 rb s3://labex-team-documents

The output is remove_bucket: labex-team-documents. Confirm storage still responds:

aws s3 ls

There are no bucket rows in this fresh workspace. AWS View shows No buckets; an Unavailable message would mean the page cannot establish the state. Your local source and downloaded documents remain available for review.

Summary

You organized two departments' documents with complete object keys and prefixes, assigned standard and custom metadata, inspected properties through the S3 API, and compared downloaded bytes with their source. You also previewed a prefix-scoped deletion, preserved an unrelated document, and cleaned up the remaining lab-owned resources.

Use prefixes for predictable name-based grouping and metadata to describe objects. Choose deletion scopes deliberately and verify the resources that should remain.