Acceldata
Pulse

Modify Storage Retention

Pulse uses multiple storage layers for different purposes, and all database layers store data under the following path:

$AcceloHome/data
The following sections describe how to override the default retention settings based on the available storage capacity on the server.

MongoDB Retention Configuration

Default Retention

  • MongoDB retention: 15 days
  • MongoDB HDFS reports retention: 15 days

Configuration

  • The Gauntlet configuration file acts as the source of truth for Gauntlet and is synchronized with the database.
  • To update the configuration, use one of the following methods:
    • Delete the existing Gauntlet configuration file and regenerate it using values from accelo.yml.
    • Manually edit the existing Gauntlet configuration file and push the updated changes to the database.

Step 1: Configure Retention

Run the following command on the Pulse node:

accelo config retention
You will be prompted to configure retention values.
Example
INFO: Using default value for mongo-retention-days as none has been provided by user: 15
How many days of data would you like to retain at Mongo DB ?: 15
INFO: Using default value for mongo-hdfs-retention-days as none has been provided by user: 15
How many days of data would you like to retain at Mongo DB for HDFS reports ?: 15
INFO: Using default value for tsdb-retention-days as none has been provided by user: 31
INFO: Using default value for snap-cleanup-freq-hours as none has been provided by user: 0
How many days of data would you like to retain at TSDB ?: 31
How often should Mongo DB clean up & compaction run, provide a comma separated string of hours (valid values are [0,23] (Ex. 8,12,15,18)?: 0
Is Mongo DB sharded [y/n] ?: n
INFO: Updating accelo.yml with gauntlet config info
INFO: Updated accelo.yml
This updates the following configuration file:
$AcceloHome/config/accelo.yml

Step 2: Generate/Edit Data Retention Configuration

If the Gauntlet configuration file is not present inside this directory:

$AcceloHome/config/gauntlet
Then, run the following command.
accelo admin database push-config -s
This generates the MongoDB retention configuration file:
$AcceloHome/config/gauntlet/gauntlet_mongo_<CLUSTER_NAME>.yml
If you want to update the retention values, you can update them manually in the above file.

Step 3: Verify MongoDB Retention Configuration

Open the generated file and verify the retention configuration.

vi $AcceloHome/config/gauntlet/gauntlet_mongo_<CLUSTERNAME>.yml
Example
uri: mongodb://accel:ACCELUSER_01082018@ad-db:27017
database: <CLUSTERNAME>
encrypted: false
snap_mongo_cleanup_frequency_in_hours: 0
is_sharded: false
collection_shard_key_map:
yarn_tez_queries: startTime
yarn_tez_queries_details: startTime
yarn_mr_queries: starttime
yarn_yarnapps: startedTime
yarn_spark_executor_summary: addTime
yarn_spark_jobs: submissionTime
yarn_spark_stages: submissionTime
yarn_sparkappcpu: start
yarn_sparkquery: submissionTime
yarn_sparkquerycontext: start
impala_queries: start_time
impala_query_details: start_time
service_recommendations_evals: createdAt
keep:
- description: "remove and compact collections mentioned in the list with query field time"
collections:
- hbase_regionserver_tasks
filters:
- filter: age
query_field: time
format: unix_nano_milli
unit: days
range: 15
order: DESC
multi: true
compact: true
exit_on_error: false

Step 4: Update Retention Period

Modify the following parameter based on your retention requirement:

Retention Range

What it does ?

range: 5

Retains data for 5 days.

range: 10

Retains data for 10 days.

You can also configure separate retention rules for specific collections.

Example

- description: "remove and compact collections mentioned in the list with query field addTime"
collections:
- yarn_spark_executor_summary
filters:
- filter: age
query_field: addTime
format: unix_nano_milli
unit: days
range: 15
order: DESC
multi: true
compact: true
exit_on_error: false

Victoria DB (TSDB) Retention Configuration

Default Retention

  • Data in the TSDB (Victoria) database is retained for 31 days by default.

Step 1: Configure Retention

Run the following command on the Pulse node:

accelo config retention
You will be prompted to configure retention values.
Example
INFO: Using default value for mongo-retention-days as none has been provided by user: 15
How many days of data would you like to retain at Mongo DB ?: 15
INFO: Using default value for mongo-hdfs-retention-days as none has been provided by user: 15
How many days of data would you like to retain at Mongo DB for HDFS reports ?: 15
INFO: Using default value for tsdb-retention-days as none has been provided by user: 31
INFO: Using default value for snap-cleanup-freq-hours as none has been provided by user: 0
How many days of data would you like to retain at TSDB ?: 31
How often should Mongo DB clean up & compaction run, provide a comma separated string of hours (valid values are [0,23] (Ex. 8,12,15,18)?: 0
Is Mongo DB sharded [y/n] ?: n
INFO: Updating accelo.yml with gauntlet config info
INFO: Updated accelo.yml
This updates the following configuration file:
$AcceloHome/config/accelo.yml
Step 2: Apply Updated Configuration
After making changes, push the configuration again:
accelo admin database push-config
Step 3: Restart the services so that the new configuration is applied
accelo restart all -d

Elasticsearch Retention Configuration

Default Retention

  • By default, Elasticsearch retains the last 5 indices.
  • You can configure Elasticsearch retention using the supported Gauntlet retention configuration.

Step 1: Push Database Configuration

If the Gauntlet configuration file is not present inside this directory:

$AcceloHome/config/gauntlet
Then, run the following command.
accelo admin database push-config -s
This generates the Elasticsearch retention configuration file:
$AcceloHome/config/gauntlet/gauntlet_elastic_<CLUSTER_NAME>.yml

Step 2: Verify Elasticsearch Retention Configuration

Notes

  • By default, Elasticsearch retains the last 5 indices.
  • Update the address field only when LogSearch is deployed on a different node.
  • No configuration changes are required if Elasticsearch is running on the core node.

Example

version: 1
elastic_servers:
- version: v8
address: "http://ad-elastic:9200"
enable_tls: false
client_certificate_path: ""
client_key_path: ""
client_ca_cert: ""
actions:
- priority: 1
keep:
- description: Keep only upto last 5 indices and delete the rest
filters:
- filter: age
format: yyyy.mm.dd
unit: days
range: 5
order: DESC
- description: Keep only upto last 2 indices for 'hdfs_audit' indices and delete the rest
filters:
- filter: regex
value: '.*hdfs-audit.*'
delete_on_match: true
- filter: age
format: yyyy.mm.dd
unit: days
range: 2
order: DESC

Step 3: Update Retention Value

Modify the following parameter based on your retention requirement:

Retention Range Values

What it does ?

range: 5

Retains the last 5 indices.

range: 10

Retains the last 10 indices.

You can also configure separate retention rules for specific index patterns.

Example:

version: 1
elastic_servers:
- version: v8
address: "http://ad-elastic:9200"
enable_tls: false
client_certificate_path: ""
client_key_path: ""
client_ca_cert: ""
actions:
- priority: 1
keep:
- description: Keep only upto last 5 indices and delete the rest
filters:
- filter: age
format: yyyy.mm.dd
unit: days
range: 5
order: DESC
- description: Keep only upto last 2 indices for 'hdfs_audit' indices and delete the rest
filters:
- filter: regex
value: '.*hdfs-audit.*'
delete_on_match: true
- filter: age
format: yyyy.mm.dd
unit: days
range: 2
order: DESC

Step 4: Apply Updated Configuration

After making changes, push the configuration again:

accelo admin database push-config