Course
Big Data using Apache Suite
15 hours 53 minutes
Credits: Optional Learning
Description
In this course, learners will develop hands-on skills in three foundational big data technologies – Hadoop, Apache HBase, and Apache Spark. Learners will explore the Hadoop Distributed File System, cluster management, cloud deployment on Amazon EMR, security with Ranger, and distribution and maintenance practices. They will then work with Apache HBase to install and configure the database, understand its architecture and data modelling principles, and access and manage data through the shell and Java Client API, including filters, snapshots, backups, and MapReduce integration. The course concludes with a comprehensive introduction to Apache Spark, covering its architecture and execution model, RDDs, DataFrames, Spark SQL, and practical data transformation, cleaning, and preparation techniques including joins, aggregations, and window functions.
What Students Will Learn
- Hadoop Distributed File System
- Clusters
- Hadoop on Amazon EMR
- Hadoop Ranger
- Maintenance & Distributions
- HBase Installation
- Apache HBase Fundamentals: Installation, Architecture, and Data Modeling
- Access Data through the Shell
- Access Data through the Java Client API
- Filters and Administration
- Snapshots, Backups, and MapReduce
- Introduction to Apache Spark
- Spark Architecture and Execution Model
- Setting Up and Navigating a Spark Environment
- Working with RDDs – Spark's Core Data Abstraction
- DataFrames and the DataFrame API
- Querying Data with Spark SQL
- Data Transformation Using Spark – Joins, Aggregations, and Window Functions
- Data Cleaning and Preparation with Spark
Overall Learning Outcomes
- Describe the Hadoop Distributed File System architecture and manage Hadoop clusters effectively
- Deploy and configure Hadoop on Amazon EMR and apply security controls using Hadoop Ranger
- Maintain and administer Hadoop distributions in production environments
- Install and configure Apache HBase and explain its architecture and data modelling principles
- Access and manipulate HBase data using the shell and Java Client API
- Apply filters, snapshots, backups, and MapReduce to manage and process HBase data
- Explain Apache Spark’s architecture, execution model, and core data abstractions including RDDs and DataFrames
- Set up and navigate a Spark environment for distributed data processing
- Query and analyse data using Spark SQL
- Transform, join, aggregate, and apply window functions to large datasets using Spark
- Clean and prepare data for downstream analytics using Apache Spark

