ODP
Search documentation…
⌘K
Documentation
Support Matrix
KB Articles
Version
ODP
ODP 3.3.6.4-1
ODP 3.3.6.3-1
ODP 3.3.6.2-102
ODP 3.3.6.2-1
ODP 3.3.6.1-1
ODP 3.3.6.0-1
ODP 3.2.3.7-3
ODP 3.2.3.7-2
ODP 3.2.3.6-3
ODP 3.2.3.6-2
ODP 3.2.3.5-3
ODP 3.2.3.5-2
ODP 3.2.3.4-3
ODP 3.2.3.4-2
ODP 3.2.3.3-3
ODP 3.2.3.3-2
ODP 3.2.3.2-3
ODP 3.2.3.2-204
ODP 3.2.3.2-203
ODP 3.2.3.2-2
ODP 3.2.3.1-3
ODP 3.2.3.1-2
ODP 3.2.3.0-3
ODP 3.2.3.0-2
Esc
Release Notes
Acceldata ODP 3.3.6.4-1 Release Notes
Supported Apache Component Versions
New Features
Enhancements
Bug Fixes
Fixed CVEs
Apache Release Notes
Known Limitations
What is ODP
Introduction to Open Source Data Platform (ODP)
Search Acceldata Documentation
Installation
Getting Started
Apache Ambari Installation and Cluster Setup
Pre-installation Preparation
Software and Memory Requirements
Install Python 3.11 in Ubuntu 20
Maximum Open Files Configuration
Data Collection and Environment Preparation
Setting up the Environment
Configuring Password-less SSH
Enable NTP on Cluster and Browser Host
Edit the Host File and Set the Hostname
Edit the Network Configuration File
Configuring iptables
Disable SELinux, PackageKit, and check the umask Value
Obfuscating LDAP Bind Password in Ambari
Using an Existing Database or Installing a Default One
Using a Local Repository
Deploying ODP In Production Data Centers With Firewalls
Accessing Acceldata Repositories
Ambari Repositories
ODP Stack Repositories
Utility Repositories
Installing Ambari
Installing the Ambari Server
Set up the Ambari Server
Installing the Ambari Agents Manually
Cluster Deployment and Configuration
Start the Ambari Server
Log in to Apache Ambari
Launch the Cluster Install Wizard
Cluster Naming and Version Selection
Configuring Ambari with RedHat Satellite or Spacewalk
Installation Options
Host Confirmation
Service Selection
Master, Slave, and Client node Assignment
Service Customization
Review Install, Start, or Test
Finalization and Completion
Guide for Management Packs
Adding a New Node to the ODP Cluster
Using Oracle Database with ODP
Using Ambari with Oracle
Using Hive with Oracle
Using Oozie with Oracle
Using Ranger and RangerKMS with Oracle
Using Hue with Oracle
Using Schema Registry with Oracle
Component User guide and Installation Instructions
Install Oozie
Airflow
Single Node Installation
Multi Node Installation
Advanced Configurations
Installing Custom Packages in Apache Airflow
Airflow Log Clean-up and DAG Management
Apache Airflow Logging Guide
Running Apache Airflow Commands via CLI
Install ClickHouse
Understand ClickHouse
Explore the Architecture
Follow Best Practices
Install ClickHouse
Manage Configuration Files
Apply Security and Governance
Connect Integrations
Use Interfaces and Tools
Druid
Enabling Kerberos in Druid
Enabling LDAP on Druid
Enabling SSL on Druid
Obfuscating LDAP Bind Password for Druid
Installing Flink and Usage
Installing Impala
JupyterHub
Jupyter Prerequisites
Installing JupyterHub via Mpack
Handling HDFS and YARN Permissions
Configuring YarnSpawner and HDFSCM
Debugging Issues with YarnSpawner and HDFSCM
Enabling SSL for JupyterHub
Configuring Databases for JupyterHub
Jupyter Authentication
Documentation for Spark Notebook Examples
Error in Submitting YARN Job
Custom Packages Installation in JupyterHub Notebook
Tensorflow with Jupyterhub
Ray with JupyterHub
Pytorch on Jupyterhub
Streamlit on Jupyterhub
Jupyterhub with Multi Spark and Additional Kernals
Kafka
Kafka Connect
Kafka MirrorMaker2
Cruise Control
Configure Kafka Tiered-Storage with S3
Installing Kafka 3
Kafka3 Connect
Kafka3 MirrorMaker2
Kafka3 Cruise Control
Zookeeper to Kraft Migration
Deploy Standalone Kafka
Deploy Standalone Cruise Control
Installing Knox
Configure Okta SSO for Apache Knox
Kudu
Install Kudu
Install Kudu using Ambari Mpack
Uninstall Kudu using Ambari Mpack
Configure and Secure Kudu
Configure Kudu Flags, SSL, and Superuser ACLs
Secure Kudu with Kerberos and Ranger
Enable Audit Support for Kudu
Enable Encryption at Rest for Kudu
Administer Kudu
Access and Use Kudu Tables
Check Kudu Cluster Health with ksck
Back Up and Restore Kudu Tables
Add a Kudu Master
Perform a Rolling Restart
Apply Final Configuration Updates
Additional Administration Considerations for Apache Kudu
Verify Kudu Status After Migration
Remove Kudu Masters
Remove Kudu Masters (Ambari Multi-Master)
Important Considerations for Master Removal and Recovery
Known Limitations in Apache Kudu
Modern Open Table Format integration with ODP
Hudi
Iceberg
Delta Lake
Deploy the Delta-Hive Connector
Installing NiFi
NiFi Upgrade Toolkit: High-Level Design for Upgrading ODP NiFi 1.x to 2.7.2
NiFi Upgrade Standalone Toolkit — Usage Guide
3-Node NiFi 2.7.2 Standalone Cluster Setup Guide
Ozone
Ozone Configurations
Ozone Installation
Ozone Command Line Interface
Working with Ozone File System
Using Ozone S3 Gateway
Install Ozone 2
Ozone2 Storage Elements
Ozone Environment Details
Configure Ozone
Install Ozone2
Ozone2 Command Line Interface
Working with Ozone File System
Using Ozone2 S3 Gateway
Ozone2 Known Issues / Limitations
Pinot
Introduction to Apache Pinot
Apache Pinot Sizing Guide
Install Pinot via Mpack
Install Knox
Batch Import using Local File
Streaming Ingestion Example with Kafka 2/Kafka 3
Configure Schema and Tables in a Kerberized Environment
Apache Pinot Batch Import using HDFS
Operational Workflow of Pinot
Leverage HDFS as Deep Storage for Real-Time Tables
Pinot Known Limitations
Schema Registry
Apache Spark 4.1.1
Connectivity
Structured Streaming
Feature Summary
Spark Reference
SQL Features
PySpark Improvements
Data Sources and Extensions
Custom Functions
Usability
Known Limitations
Trino
Introduction to Trino
Trino Capacity Planning
Install Trino via Mpack
Set up Trino SSL
Set up Trino Knox
Set up Trino Ranger (For 3.3.6.x-x or later)
Trino Supported Connectors
Configure JMX
Troubleshooting Trino
Configure Trino Client
Trino Known Limitations
Install Trino Gateway
Introduction to Trino Gateway
Prerequisites and Requirements
Managing the Trino Gateway Mpack
Installing and Configuring Trino Gateway
Trino Gateway Usage
Querying Trino Through Trino Gateway
Routing Rules
Trino Gateway LDAP Authentication
Trino Gateway Known Issues
Install MLflow
Get Started with MLflow
Manage the ML Lifecycle using MLflow
Use MLflow in Real-World Projects
Set up MySQL as MLflow's Backend Store
Set up PostgreSQL as MLflow’s Backend Store
Install MLflow using Ambari Mpack
Run MLflow in Local or Remote Mode
Store MLflow Artifacts in HDFS
Store MLflow Artifacts in S3
Install NGINX and Configure for MLflow
Configure Basic Authentication for MLflow
Integrate HBase-Spark
Run Spark Rapids
Zeppelin
Obfuscating LDAP Bind Password for Zeppelin
Ranger S3
Why Ranger-S3: Addressing Current Limitations
S3 Plugin Implementation Architecture
Ranger Prerequisites
Ranger Implementation
Ranger Limitations
Install ODP Celeborn
Introduction to Celeborn
Celeborn vs YARN External Shuffle Service
Architecture Overview
Prerequisites for Celeborn
Deploy Celeborn with Ambari Mpack
Configuration Reference
Engine Integration
Security Configuration
Configure Storage
Quota Management
Worker Tags
Monitoring and Metrics
Performance Tuning
Capacity Planning
Troubleshooting
Celeborn Appendix
Install Superset
Introduction
Install Superset
Connector Support
Superset Pinot testing - Batch Data
Superset Pinot testing - Stream Data
Superset Impala Testing
Superset Hive Testing
Superset Druid Testing
Superset Clickhouse Testing
Superset Trino Testing
Superset MySQL Testing
Superset Postgres Testing
Access the Superset UI
Administration and User Roles
Uninstall Superset
Spark Connect
Upgrade Instructions
Upgrade Path
Master Upgrade
Master Upgrade from 3.3.6.x-x to 3.3.6.4-1
Backup Ambari Configuration (Mpacks)
Prerequisites for Upgrade
Create and Update VDF File
Configure Backup for Stack Components
Stop, Delete, and Uninstall Mpack Services from Ambari
Service Checks for Components
Disable Service Auto-Start
Install ODP Packages
ODP Stack Upgrade
ODP Rolling upgrade
ODP Express Upgrade
Ambari Upgrade
Upgrade Ambari Server
Upgrade Ambari Agents
Upgrade Infra-solr
Configure JDK 17 Environment Changes
Configure the Java 17 Flags Automatically Using a Script
Configure the Java 17 Flags Manually
Hadoop
YARN
MapReduce
Tez
Hive
Oozie
RangerKMS
Infra-solr
Resume and Finalize the ODP Stack Upgrade
Master Upgrade from 3.2.3.x-x to 3.3.6.4-1
Backup Ambari Configuration (Mpacks)
Prerequisites for Upgrade
Create and Update VDF File
Configure Backup for Stack Components
Stop, Delete, and Uninstall Mpack Services from Ambari
Service Checks for Components
Disable Service Auto-Start
Install ODP Packages
Upgrade ODP using the Express Upgrade option
ODP Upgrade Halt State
Ambari Upgrade
Ambari Server Upgrade
Ambari Agents Upgrade
Infra Solr Upgrade
Configure JDK 17 Environment Changes
Configure the Java 17 Flags Manually
Hadoop
YARN
MapReduce
Druid
HBase
Hive
Tez
Oozie
Infra-solr
RangerKMS
Configure the Java 17 Flags Automatically Using a Script
Resume and Finalize the ODP Stack Upgrade
Sanity Check
Troubleshooting
Upgrade: Known Issues
FAQs About Upgrade
Upgrade Standalone Kafka
Rolling Upgrade from Kafka 2 to Kafka 3
In-Place Upgrade from Kafka 2 to Kafka 3
Sidecar (Rolling) Upgrade from Kafka 2 to Kafka 3
Downgrade Instructions
Downgrade Instructions
Reference Guide
Ambari Admin Guide
Using the Administrator Role in Ambari Web
Setting up Ambari to use an Internet proxy server
Managing Cluster Roles
Managing Versions
Managing Local Users
Ambari High Availability Setup
Administration
Administering HDFS
Managing and Monitoring a Cluster
Managing High Availability
Configuring Dynamic Management for HDFS NameNode and DataNode
Enhance Replication and Decommissioning process
Setting Hadoop Service Log Levels at Runtime
Storage
Configuring ADLS Gen2 with ODP
Configuring ODP with GCS
Configuring ODP with AWS S3 using S3A
Configure Apache Hive with S3A
Configuring Ports
Controlling ODP Services Manually
HDFS Router Federation
Set up the ODBC and JDBC Connections
Set up the Hive and Impala ODBC Connection
Set up the Hive and Impala JDBC Connection
Securing Hive with KNOX
Set up HAProxy Load Balancer for HiveServer2
Service Tuning
HDFS: Separate RPC Queues and Enable Lifeline Protocol
Erasure Coding for Data Durability
Understanding Erasure Coding Policies
Comparing Replication and Erasure Coding
Best Practices for Rack and Node Setup for EC
Prerequisites for Enabling Erasure Coding
Limitations of Erasure Coding
Using Erasure Coding for Existing Data
Using Erasure Coding for New Data
Advanced Erasure Coding Configuration
Erasure Coding CLI Commands
Erasure Coding Examples
Guidelines for Enabling Ranger Auditing on Supported Plugins
Enabling Ranger HDFS Audits
Enabling Ranger Solr Audits
Migrating Oozie Workflows to Airflow DAGs (O2A)
Convert Oozie Workflows to Airflow DAGs
Bulk Workflow Migration
Shell Example
Hive/Hive2 Example
Known Limitations (O2A)
Errors You May Encounter
Unsupported Features (O2A)
Security Guide
Authentication
Replacing Knox Self-Signed Certificate with CA Certificate
Enabling SSL for Services Using Bash Script
Enabling Kerberos in an ODP Cluster
Bash Scripts for Automated Tasks
Securing NiFi with Existing CA Certificates
Configuring SSL for the Ambari Server
Accessing a Kerberized UI Firefox
Configure Ranger Admin and Usersync with SSSD or Centrify
Enabling Cross-Cluster and Cross-Realm Kerberos Authentication for Hadoop Data Migration
Authorization
Providing Authorization with Apache Ranger
Configuring Ambari Authentication for LDAP/AD
Securing HiveServer2 with LDAP Authentication
Configuring Ranger Admin High Availability (HA)
Encryption
Enabling the ODP Ranger KMS High Availability (HA) and Troubleshooting
Disabling High Availability (HA) in ODP Ranger KMS
Copying the Encrypted Data Between Two ODP Clusters Using Ranger KMS
ODP Password Obfuscation Guide
Troubleshooting Guide
Troubleshooting ODP
Ambari Troubleshooting
HDFS NameNode Heap Estimation
Troubleshooting premature swapping
Troubleshooting the Kerberos Configuration Issues in Ambari
Troubleshooting JVM: JStack, Heap Dumps and JFR
HBase Consistency Checker 2 (HBCK2)
Introduction to HBCK2
General Usage and Invocation
HBCK 2 Commands
addFsRegionsMissingInMeta (Hbck2)
assigns (Hbck2)
bypass (Hbck2)
extraRegionsInMeta (Hbck2)
filesystem (Hbck2)
fixMeta (Hbck2)
generateMissingTableDescriptorFile (Hbck2)
recoverUnknown (Hbck2)
regionInfoMismatch (Hbck2)
replication (Hbck2)
reportMissingRegionsInMeta (Hbck2)
setRegionState (Hbck2)
setTableState (Hbck2)
scheduleRecoveries (Hbck2)
unassigns (Hbck2)
Troubleshooting ODP SSL, Keystore, and Truststore
What is SSL?
Key Concepts: Keystore and Truststore
Types of Certificates
How to Check the Certificate Type?
JKS Creation and Management
Extracting Certificates and Keys
Checking Subject Alternative Name (SAN) and Extensions
Creating a Keystore in PKCS12 Format
TLS Verification
Checking SSL Ciphers
Additional SSL Troubleshooting Commands
Troubleshooting Ambari SSL and LDAPS Setup
Services SSL and Truststore Configuration
Configure the Ranger Plugin Truststore SSL
SSL Troubleshooting for Hadoop Clusters
SSL/TLS Security Best Practices for Hadoop Clusters
Uninstall ODP
Complete Uninstallation of ODP
ODP
›
Installation
›
Launch the Cluster Install Wizard
Launch the Cluster Install Wizard
From the Ambari Welcome page, choose
Launch Install Wizard
.
Was this page helpful?
0 helpful · 0 not helpful
Yes
No
Send feedback