IEC provides comprehensive monitoring functions for you to understand system resource usage and performance in real time. Integrating Prometheus and Grafana, you can view the usage of CPU, memory, storage and network, as well as the running status of each workload through the monitoring dashboard. You can also view the historical data by specifying a time range.
Prometheus rules including alerting rules and recording rules to set an alert rule and group/filter notifications. These rules define when an alarm will be triggered, such as resource usage exceeding the threshold. Routes define how to organize, filter, and control the sending process of triggered alarms. Receivers define who receives these alert messages and how. You can set alert rules, grouping, filtering and notification methods to ensure that relevant personnel can be notified immediately via email when there is an abnormality in the system or application. Monitor, manage, process and transmit alarm information through configuration settings such as configuring rules, configuring routing rules and configuring receivers.
The Grafana dashboard visually displays resource consumption, helping you identify and adjust resource configurations as needed to avoid potential resource bottlenecks or waste.
For each workload, you can view its health status, resource requirements and performance data including resource consumption, network traffic, number of requests, etc, enabling you to respond to possible failures or abnormalities in a timely manner.
For overall cluster health, IEC provides comprehensive monitoring such node status, pods restart status, and disk space availability.