Introduction
In the Hadoop versions 2.0 and later, the new resource management pattern of YARN was introduced, which facilitates the cluster in terms of utilization, unified resource management and data sharing. Based on the foundation of building the Hadoop pseudo-distributed cluster, this section will let you learn the architecture, the working principle, configuration, and development and monitoring techniques of the YARN framework.
This lab requires a certain Java programming foundation.
The development steps provide complete Java source files. Paste each full file into your editor, then review how the client submits an application and how the ApplicationMaster requests and launches a task container.
YARN Architecture and Components
In this step, explore YARN architecture and the roles of its components.
YARN, introduced in Hadoop 0.23 as part of MapReduce 2.0 (MRv2), revolutionized the resource management and job scheduling in Hadoop clusters:
- Decomposition of JobTracker:
MRv2decomposes the functions ofJobTrackerinto separate daemons -ResourceManagerfor resource management andApplicationMasterfor job scheduling and monitoring. - Global ResourceManager: Each application has a corresponding
ApplicationMaster, which can be a MapReduce job or a DAG describing the job. - Data Calculation Framework: The
ResourceManager,Slave, andNodeManagerform a framework where the ResourceManager governs all application resources. - ResourceManager Components: The
Schedulerallocates resources based on constraints like capacity and queues, while theApplicationsManagerhandles job submissions and ApplicationMaster execution. - Resource Allocation: Resource requirements are defined using resource containers with elements like memory, CPU, disk, and network.
- NodeManager Role: NodeManager monitors container resource usage and reports to ResourceManager and Scheduler.
- ApplicationMaster Tasks: ApplicationMaster negotiates resource containers with Scheduler, tracks status, and monitors progress.
The following figure depicts the relationship:

YARN ensures API compatibility with previous versions, allowing seamless transition for running MapReduce tasks. Understanding the architecture and components of YARN is essential for efficient resource management and job scheduling in Hadoop clusters.
Starting the Hadoop Daemon
In this step, start the Hadoop daemons needed to run the YARN application.
Before learning the relevant configuration parameters and YARN application development techniques, we need to start the Hadoop daemon so that it can be used at any time.
First double-click to open the Xfce terminal on the desktop and enter the following command to switch to the hadoop user:
su - hadoop
tip: The password is 'hadoop' of the user 'hadoop'.
Once the switch is complete, you can start the Hadoop-related daemons including HDFS and YARN frameworks.
Please enter the following commands in the terminal to start the daemons:
/home/hadoop/hadoop/sbin/start-dfs.sh
/home/hadoop/hadoop/sbin/start-yarn.sh
After the boot has completed, you can choose to use the jps command to check if the associated daemons are running.
hadoop:~$ jps
3378 NodeManager
3028 SecondaryNameNode
3717 Jps
2791 DataNode
2648 NameNode
3240 ResourceManager
Preparing the Configuration File
In this step, we will learn about yarn-site.xml, which is one of the main configuration files of Hadoop, to see what settings can be made for the YARN cluster in this file.
To prevent the misuse of changes to the configuration file, it is the best to copy the Hadoop configuration file to another directory and then open it.
To do this, please enter the following command in the terminal to create a new directory for the configuration file:
mkdir /home/hadoop/hadoop_conf
Then copy the main configuration file yarn-site.xml of YARN from the installation directory to the newly created directory.
Please enter the following command in the terminal to perform the operation:
cp /home/hadoop/hadoop/etc/hadoop/yarn-site.xml /home/hadoop/hadoop_conf/yarn-site.xml
Then use vim editor to open the file to view its content:
vim /home/hadoop/hadoop_conf/yarn-site.xml
How the Configuration File Works
In this step, review the configuration parameters used by the running Hadoop cluster.
We know that there are two important roles in the YARN framework: ResourceManager and NodeManager. Therefore, each configuration item in the file is a setting of the above two components.
There are many configuration items that can be set in this file, but by default, this file does not contain any custom configuration items. For example, the file we open now has only the attribute aux-services that was specified when the pseudo-distributed Hadoop cluster was previously configured, as shown in the following figure:
hadoop:~$ cat /home/hadoop/hadoop/etc/hadoop/mapred-site.xml
...
<configuration>
<property>
<name>mapreduce.framework.name</name>
<value>yarn</value>
</property>
</configuration>
This configuration item is used to set the dependent services that need to be run on the NodeManager. The configuration value we specify is mapreduce_shuffle, indicating that the default value of the MapReduce program needs to be run on YARN.
Are the configuration items not written in it not working? Not exactly. When the configuration parameters are not explicitly specified in the file, Hadoop's YARN framework will read the default values stored in internal files. All configuration items explicitly specified in the yarn-site.xml file will override the default values, which is an effective way for the Hadoop system to adapt to different usage scenarios.
ResourceManager Configuration Items
Understanding and correctly configuring the ResourceManager settings in the yarn-site.xml file is essential for efficient resource management and job execution in a Hadoop cluster. Here is a summary of the key configuration items related to the ResourceManager:
yarn.resourcemanager.address: Exposes the address to clients for submitting applications and killing applications. Default port is 8032.yarn.resourcemanager.scheduler.address: Exposes the address to ApplicationMaster for requesting and releasing resources. Default port is 8030.yarn.resourcemanager.resource-tracker.address: Exposes the address to NodeManager for sending heartbeats and pulling tasks. Default port is 8031.yarn.resourcemanager.admin.address: Exposes the address to administrators for management commands. Default port is 8033.yarn.resourcemanager.webapp.address: WebUI address for viewing cluster information. Default port is 8088.yarn.resourcemanager.scheduler.class: Specifies the main class name of the scheduler (e.g., FIFO, CapacityScheduler, FairScheduler).- Thread Configuration:
yarn.resourcemanager.resource-tracker.client.thread-countyarn.resourcemanager.scheduler.client.thread-count
- Resource Allocation:
yarn.scheduler.minimum-allocation-mbyarn.scheduler.maximum-allocation-mbyarn.scheduler.minimum-allocation-vcoresyarn.scheduler.maximum-allocation-vcores
- NodeManager Management:
yarn.resourcemanager.nodes.exclude-pathyarn.resourcemanager.nodes.include-path
- Heartbeat Configuration:
yarn.resourcemanager.nodemanagers.heartbeat-interval-ms
Configuring these parameters allows fine-tuning of ResourceManager behavior, resource allocation, thread handling, NodeManager management, and heartbeat intervals in a Hadoop cluster. Understanding these configuration items helps prevent issues and ensures smooth operation of the cluster.
NodeManager Configuration Items
Configuring the NodeManager settings in the yarn-site.xml file is crucial for managing resources and tasks efficiently within a Hadoop cluster. Here is a summary of the key configuration items related to the NodeManager:
yarn.nodemanager.resource.memory-mb: Specifies the total physical memory available to the NodeManager. This value remains constant throughout YARN runtime.yarn.nodemanager.vmem-pmem-ratio: Sets the ratio of virtual memory to physical memory allocation. Default ratio is2.1.yarn.nodemanager.resource.cpu-vcores: Defines the total number of virtual CPUs available for the NodeManager. Default value is8.yarn.nodemanager.local-dirs: Path for storing intermediate results on the NodeManager, allowing configuration of multiple directories.yarn.nodemanager.log-dirs: Path to the log directory of the NodeManager, supporting configuration of multiple directories.yarn.nodemanager.log.retain-seconds: Maximum retention time for NodeManager logs, default is 10800 seconds (3 hours).
Configuring these parameters enables fine-tuning of resource allocation, memory management, directory paths, and log retention settings for optimal performance and resource utilization by the NodeManager in a Hadoop cluster. Understanding these configuration items helps ensure smooth operation and efficient task execution within the cluster.
Configuration Item Query and Default References
To explore all configuration items available in YARN and other common Hadoop components, you can refer to the default configuration files provided by Apache Hadoop. Here are the links to access the default configurations:
YARN Configuration Items:
Common Configuration Files:
- core-default.xml (core-site.xml)
- hdfs-default.xml (hdfs-site.xml)
- mapred-default.xml (mapred-site.xml)
Exploring these default configurations provides detailed descriptions of each configuration item and their purposes, helping you understand the role of each parameter in Hadoop architecture design.
After reviewing the configurations, you can close the vim editor to conclude your exploration of Hadoop configuration settings.
Creating Project Directories and Files
In this step, create the source files for our application. We will build a complete, minimal YARN application: a client submits the ApplicationMaster, which requests one task container to print a greeting.
First create a project directory. Please enter the following command in the terminal to perform directory creation:
mkdir /home/hadoop/yarn_app
Then create two source code files in the project separately.
The first one is Client.java. Please use the touch command in the terminal to create the file:
touch /home/hadoop/yarn_app/Client.java
Then create the ApplicationMaster.java file:
touch /home/hadoop/yarn_app/ApplicationMaster.java
hadoop:~$ tree /home/hadoop/yarn_app/
/home/hadoop/yarn_app/
├── ApplicationMaster.java
└── Client.java
0 directories, 2 files
Writing Client Code
In this step, write the complete client that submits our application. Continue as the hadoop user. Open the source file:
vim /home/hadoop/yarn_app/Client.java
Replace the entire file with the code below, including the package declaration and imports. In Vim, type :set paste and press Enter before pressing i to enter insert mode. Paste the complete file, wait for all lines to appear, then press Esc and type :wq to save and exit.
package com.labex.yarn.app;
import java.util.Collections;
import java.util.HashMap;
import java.util.Map;
import org.apache.hadoop.fs.FileStatus;
import org.apache.hadoop.fs.FileSystem;
import org.apache.hadoop.fs.Path;
import org.apache.hadoop.yarn.api.ApplicationConstants;
import org.apache.hadoop.yarn.api.records.*;
import org.apache.hadoop.yarn.client.api.YarnClient;
import org.apache.hadoop.yarn.client.api.YarnClientApplication;
import org.apache.hadoop.yarn.conf.YarnConfiguration;
import org.apache.hadoop.yarn.util.ConverterUtils;
public class Client {
public static void main(String[] args) throws Exception {
if (args.length != 1) {
throw new IllegalArgumentException("Usage: Client /absolute/path/to/yarn-app.jar");
}
YarnConfiguration conf = new YarnConfiguration();
YarnClient client = YarnClient.createYarnClient();
client.init(conf);
client.start();
try {
YarnClientApplication application = client.createApplication();
ApplicationSubmissionContext context = application.getApplicationSubmissionContext();
ApplicationId id = context.getApplicationId();
FileSystem fs = FileSystem.get(conf);
Path destination = new Path(fs.getHomeDirectory(), "yarn-app/" + id + "/app.jar");
fs.mkdirs(destination.getParent());
fs.copyFromLocalFile(new Path(args[0]), destination);
FileStatus status = fs.getFileStatus(destination);
LocalResource jar = LocalResource.newInstance(
ConverterUtils.getYarnUrlFromPath(fs.makeQualified(destination)),
LocalResourceType.FILE, LocalResourceVisibility.APPLICATION,
status.getLen(), status.getModificationTime());
Map<String, LocalResource> resources = new HashMap<>();
resources.put("app.jar", jar);
Map<String, String> environment = new HashMap<>();
// The single-node lab uses the same Hadoop installation in every container.
environment.put("CLASSPATH", "./app.jar:" + System.getProperty("java.class.path"));
String command = ApplicationConstants.Environment.JAVA_HOME.$$()
+ "/bin/java -Xmx128m com.labex.yarn.app.ApplicationMaster"
+ " 1>" + ApplicationConstants.LOG_DIR_EXPANSION_VAR + "/stdout"
+ " 2>" + ApplicationConstants.LOG_DIR_EXPANSION_VAR + "/stderr";
ContainerLaunchContext launch = ContainerLaunchContext.newInstance(
resources, environment, Collections.singletonList(command), null, null, null);
context.setApplicationName("LabEx YARN Hello");
context.setQueue("default");
context.setResource(Resource.newInstance(256, 1));
context.setAMContainerSpec(launch);
client.submitApplication(context);
System.out.println("Application ID: " + id);
long deadline = System.currentTimeMillis() + 180000;
while (System.currentTimeMillis() < deadline) {
ApplicationReport report = client.getApplicationReport(id);
YarnApplicationState state = report.getYarnApplicationState();
if (state == YarnApplicationState.FINISHED
|| state == YarnApplicationState.FAILED
|| state == YarnApplicationState.KILLED) {
System.out.println("Final status: " + report.getFinalApplicationStatus());
if (report.getFinalApplicationStatus() != FinalApplicationStatus.SUCCEEDED) {
throw new IllegalStateException(report.getDiagnostics());
}
return;
}
Thread.sleep(1000);
}
client.killApplication(id);
throw new IllegalStateException("Application timed out after three minutes");
} finally {
client.stop();
}
}
}
YarnClient connects to the ResourceManager and obtains an application ID. The client copies the JAR into HDFS and declares it as the local resource app.jar, so YARN can place it in the ApplicationMaster's working directory. The launch context supplies the classpath and Java command; the submission context specifies the queue and container resources.
The classpath uses the installed Hadoop libraries shared by the containers in this single-node lab. No Maven or Gradle download is required. This example targets the lab's non-Kerberos cluster.
After submission, the client polls the application report until a terminal state is reached and checks the final status. killApplication is used only if the application exceeds the three-minute timeout. We will compile and execute both classes in the final step.
Writing ApplicationMaster Code
In this step, write the complete ApplicationMaster. It registers with the ResourceManager, requests one container, launches a greeting through the NodeManager, and reports the task's result.
Open the file as the hadoop user:
vim /home/hadoop/yarn_app/ApplicationMaster.java
Replace the entire file with the code below. Type :set paste and press Enter before pressing i. Paste the complete file, wait for all lines to appear, then press Esc and type :wq to save and exit.
package com.labex.yarn.app;
import java.util.Collections;
import org.apache.hadoop.yarn.api.ApplicationConstants;
import org.apache.hadoop.yarn.api.protocolrecords.AllocateResponse;
import org.apache.hadoop.yarn.api.records.*;
import org.apache.hadoop.yarn.client.api.AMRMClient;
import org.apache.hadoop.yarn.client.api.NMClient;
import org.apache.hadoop.yarn.conf.YarnConfiguration;
public class ApplicationMaster {
public static void main(String[] args) throws Exception {
YarnConfiguration conf = new YarnConfiguration();
AMRMClient<AMRMClient.ContainerRequest> rm = AMRMClient.createAMRMClient();
NMClient nm = NMClient.createNMClient();
rm.init(conf);
nm.init(conf);
rm.start();
nm.start();
try {
rm.registerApplicationMaster("", 0, "");
AMRMClient.ContainerRequest request = new AMRMClient.ContainerRequest(
Resource.newInstance(256, 1), null, null, Priority.newInstance(0));
rm.addContainerRequest(request);
boolean launched = false;
long deadline = System.currentTimeMillis() + 120000;
while (System.currentTimeMillis() < deadline) {
// Each allocate call also sends a heartbeat to the ResourceManager.
AllocateResponse response = rm.allocate(launched ? 0.5f : 0.0f);
for (Container container : response.getAllocatedContainers()) {
if (launched) {
rm.releaseAssignedContainer(container.getId());
continue;
}
rm.removeContainerRequest(request);
String command = "/bin/echo Hello-from-YARN"
+ " 1>" + ApplicationConstants.LOG_DIR_EXPANSION_VAR + "/stdout"
+ " 2>" + ApplicationConstants.LOG_DIR_EXPANSION_VAR + "/stderr";
ContainerLaunchContext launch = ContainerLaunchContext.newInstance(
Collections.emptyMap(), Collections.emptyMap(),
Collections.singletonList(command), null, null, null);
nm.startContainer(container, launch);
launched = true;
}
for (ContainerStatus status : response.getCompletedContainersStatuses()) {
if (status.getExitStatus() != 0) {
throw new IllegalStateException(status.getDiagnostics());
}
rm.unregisterApplicationMaster(FinalApplicationStatus.SUCCEEDED,
"Hello container completed", "");
return;
}
Thread.sleep(1000);
}
rm.unregisterApplicationMaster(FinalApplicationStatus.FAILED,
"No successful container within two minutes", "");
throw new IllegalStateException("Container timed out");
} finally {
nm.stop();
rm.stop();
}
}
}
AMRMClient handles registration and resource requests. Each call to allocate sends a heartbeat and returns newly allocated containers and completed-container statuses. NMClient launches the command in the allocated container. This synchronous polling loop keeps the example self-contained without undefined callback classes or helper methods.
The task prints Hello-from-YARN to its container log. A zero exit status causes the ApplicationMaster to unregister with SUCCEEDED. The client then reports that final status. The ResourceManager can round the requested 256 MB up to its minimum allocation; the loop continues sending heartbeats while waiting.
The Process of Application Launching
In this step, compile and run the two Java classes you wrote, then inspect their application in the ResourceManager web interface. Continue in the terminal as the hadoop user.
Compiling and Launching the Application
Change to the source directory:
cd /home/hadoop/yarn_app
Create the output directory:
mkdir -p classes
Compile both complete source files against the Hadoop libraries already installed in the VM. --release 8 produces classes compatible with Hadoop's Java 8 runtime, even though the default compiler is newer:
javac --release 8 -cp "$(/home/hadoop/hadoop/bin/hadoop classpath --glob)" -d classes Client.java ApplicationMaster.java
A deprecated-API note is harmless; there should be no compilation errors. Package the compiled classes:
jar cf yarn-app.jar -C classes .
Submit your client and save its output. The final argument is the JAR that the client uploads to HDFS for the ApplicationMaster:
set -o pipefail
/home/hadoop/hadoop/bin/yarn jar /home/hadoop/yarn_app/yarn-app.jar com.labex.yarn.app.Client /home/hadoop/yarn_app/yarn-app.jar | tee /home/hadoop/yarn_app/application.log
Allow the application to finish. Among the Hadoop log messages, expect:
Application ID: application_<timestamp>_<sequence>
Final status: SUCCEEDED
The ID varies each run. The greeting is written to the task container's log, while the client prints the application ID and final result. This run exercises your own client and ApplicationMaster rather than a separate prebuilt MapReduce example.
Viewing Application Execution Results
Open Firefox on the desktop and visit:
http://localhost:8088
Find LabEx YARN Hello in the application list, using the application ID printed in your terminal. Its state should be FINISHED and its final status should be SUCCEEDED. Click the application ID to inspect its details. The ResourceManager tracks the application; the ApplicationMaster manages its task container, and the NodeManager executes that container.
Summary
Based on the completion of the Hadoop pseudo-distributed cluster, this lab continues to teach us the architecture, working principle, configuration, and development and monitoring techniques of the YARN framework. A lot of code and configuration files are given in the course, so please read them carefully.



