Skip to main content
Version: 6.1

Cluster Health Dashboard

Article Description​

The dashboard is used to analyze the state of a Smart Monitor cluster over a selected period. It shows the cluster status, changes in the number of nodes, shard status, node resource metrics, and related log messages.

The panels help identify changes in cluster state and determine the direction of further investigation.

Dashboard Composition​

Global and Local Filters​

The dashboard provides the following controls:

  • Period sets the time range for viewing data
  • Severity limits the message log by severity level
  • Search by message helps find records by message text

Color Indicators​

Color indicators help quickly identify metrics that require investigation. Use them as a guide for an initial assessment, not as independent proof of a problem.

If a color zone indicates a possible resource shortage, check the metric trend and compare it with the cluster status, shard state, and log messages.

Dashboard Metrics​

Top Status Indicators​

Example of the cluster status and active node indicators

PanelWhat it showsWhen to checkNormal-operation indicatorsWhen additional investigation is required
StatusThe overall state of shard allocation in the cluster.At the beginning of troubleshooting, when data is unavailable, or when search and/or write errors occur.The state matches the expected shard placement scheme.Check a transition to yellow if it is not related to known maintenance or persists longer than expected. red requires identifying unassigned primary shards and assessing the availability of the affected data.
Active NodesThe number of nodes participating in cluster operation.When a node loss, network problem, or service restart is suspected.The number of nodes matches the approved architecture or a known maintenance period.An unexpected decrease or frequent changes in the number of nodes should be compared with the cluster status and logs.

Shard State and Distribution​

Example of shard status charts

PanelWhat it showsWhen to checkNormal-operation indicatorsWhen additional investigation is required
Active Shards PercentageThe proportion of shards available for search.When the cluster status changes, after recovery from a failure, or when checking the completeness of shard allocation.The metric returns to the expected level after shards are recovered or relocated.Investigate a decrease if it persists, is accompanied by an increase in UNASSIGNED, or affects primary shards. The percentage alone does not show which data is affected.
Shard StatusesThe distribution of shards by state: active, initializing, relocating, and unassigned.When you need to determine whether recovery or relocation is in progress or whether there is a shard placement problem.Initializing and relocating shards eventually become active; the number of unassigned shards does not increase.Persistent INITIALIZING or RELOCATING states, as well as the appearance of UNASSIGNED, should be checked by replica type, allocation failure reason, and available resources.

Node Summary Table​

Example of the node summary table

PanelWhat it showsWhen to checkNormal-operation indicatorsWhen additional investigation is required
Node Summary TableNode resource metrics: segments, CPU, memory, storage, HTTP connections, file descriptors, heap, and threads.When you need to determine which node requires further investigation.Resource metrics are comparable between nodes with the same role and remain within operational limits.Investigate a node that stands out if the difference persists and coincides with a change in status, shard allocation, latency, or the number of errors.

Log Panels​

Example of log panels

PanelWhat it showsWhen to check
Important Messages LogCritical and warning cluster messages.After a status change, node loss, the appearance of unassigned shards, or write and/or search errors.
Cluster Messages LogDetailed cluster messages for reconstructing the event timeline.When you need to move from a general symptom to specific messages and the time they appeared.
Error AggregationRepeated messages and groups of similar errors.When you need to assess the scale of a problem and find the most frequent messages for the selected period.

Troubleshooting Examples​

Where to Find Details​

Symptom or deviationWhere to find details
Node loss or a decrease in the number of active nodesNode Resource Monitoring, Node JVM Monitoring
Unassigned shards or a change in cluster statusThe current dashboard, Node Resource Monitoring
Disk usage or shard placement problemsNode Resource Monitoring, Node Indexing Performance Monitoring
Increased search load or search errorsNode Query Performance Monitoring, Cluster Query Count by Type
Write errors or indexing degradationNode Indexing Performance Monitoring
Data collection problemsLogstash Monitoring, Logstash Node JVM Monitoring
Repeated errors in logsThe current dashboard: log panels and error aggregation

Initial Assessment of Cluster State​

Use the metric tables to determine which metric no longer meets the normal-operation indicators or meets the criteria for additional investigation. Record when the change occurred and compare values for the same time interval.

Determine the scope of the deviation: the entire cluster, an individual node, or shards of a specific index. If the cluster status or composition changes, proceed to the next section; for unassigned shards, proceed to shard allocation analysis; for resource deviations, proceed to the investigation of resource-related causes.

Analyzing Changes in Cluster Status and Composition​

First, compare the status change with the number of active nodes and events in the logs for the same period.

The yellow status means that all primary shards are available, but some replicas are not allocated. Determine why the cluster cannot allocate the required number of replicas.

The red status means that at least one primary shard is unassigned. Some data may be unavailable, and search may return incomplete results. Determine whether the change is related to node loss, insufficient resources, allocation rules, or recovery errors.

If the number of active nodes has decreased, start by checking the availability of the affected node. If the node has returned but the shards remain unassigned, proceed to analyzing the allocation failure reasons.

Analyzing Shard Allocation​

To check shard state and distribution, run a GET request in the Dev Tools console:

GET /_cat/shards?v

In the output, identify the index, shard number, replica type (p — primary shard, r — replica), state, and node where it is allocated. For a shard in the UNASSIGNED state, specify its index, number, and replica type in the body of the GET request:

GET /_cluster/allocation/explain
{
"index": "<index_name>",
"shard": 0,
"primary": true
}

The primary value is true for a primary shard and false for a replica. Without a request body, the request returns an explanation for the first unassigned shard it finds.

After identifying the affected index or shard, find the nodes in the summary table whose resource metrics changed during the same period. Compare these changes with the cluster status and shard state.

Based on the type of deviation identified, select a detailed dashboard in the Where to Find Details table and continue the investigation there.

Checking Events and Errors​

After identifying the deviation and time interval, open the Important Messages Log, followed by the Cluster Messages Log. Reconstruct the timeline: when the status changed, which nodes were affected, and which messages appeared around the change.

Then use Error Aggregation to assess how often the messages recur. Compare repeated errors with the shard state and node resources: repetition alone does not explain the source of the problem.

Typical Scenarios​

Node Exclusion​

This scenario occurs when a node loses network connectivity with the other members or stops responding to availability checks. The cluster manager excludes it from the cluster, and the shard replicas that were located on it become unavailable.

If replicas are available, OpenSearch can promote a replica to a primary shard and recover the missing replicas on other nodes. If there are no suitable copies or sufficient resources for allocation, some shards remain in the UNASSIGNED state.

The dashboard shows this as follows:

  • Status changes to yellow or red
  • the Active Nodes value decreases
  • Active Shards Percentage decreases
  • the number of unassigned shards changes on the Shard Statuses chart

Consider a two-node cluster: the first node has the master + data roles, and the second node has the data role. The number of replicas is one. The node with the data role was excluded:

Example of node exclusion

In the example, the status changed to yellow, the Active Nodes value decreased to 1, the active shards percentage decreased to approximately 70%, and some shards entered the UNASSIGNED state.

Disk Space Exhaustion​

During long-term operation without configured ISM policies or with a sharp increase in incoming data volume, the available space on Smart Monitor Data Storage nodes may be exhausted. The consequences depend on the disk threshold settings and on whether primary shards have become unavailable.

The dashboard shows this as follows:

  • Status may change to yellow or red
  • in the summary table, the Storage, % value for one or more nodes approaches 100%
  • the number of unassigned shards changes on the Shard Statuses chart

Consider a two-node cluster: the first node has the master + data roles, and the second node has the data role. The disk on the node with the data role became 100% full:

Example of a full disk

In the example, both nodes remain active, but the Storage, % value of the affected node reaches 99% in the summary table, some shards enter the UNASSIGNED state, and the status changes to red. In this case, one or more primary shards became unavailable; a full disk alone is not necessarily the cause of a red status.

Note

In Smart Monitor, disk allocation thresholds are disabled by default. If they were not enabled when the environment was configured, the cluster continues to use disk space without a protective response to the high watermark and flood_stage thresholds. When free space is exhausted, write, recovery, and shard allocation errors, node instability, and unassigned shards may occur.

For instructions on configuring Disk Watermark, see the in the corresponding article.

Note

If shards remain in the UNASSIGNED state after space has been freed, first resolve the allocation failure reason.

To manually retry shard allocation, run the following POST request:

POST /_cluster/reroute?retry_failed=true