mirror of
https://github.com/storytold/terraform-aws-observability-accelerator.git
synced 2026-10-09 00:09:43 +00:00
caeebe0e1a
* Made the cluster variable visible in dashboards * Made the repo reusable with multiple EKS clusters * Relabled the input variable from create_grafana_data_source to create_prometheus_data_source
14 KiB
14 KiB
Infrastructure monitoring
This module provides EKS cluster monitoring with the following resources:
- AWS Distro For OpenTelemetry Operator and Collector for Metrics and Traces
- Logs with AWS for FluentBit
- AWS Managed Grafana Dashboard and data source
- Alerts and recording rules with AWS Managed Service for Prometheus
This module makes use of the open source kube-prometheus-stack
Requirements
| Name | Version |
|---|---|
| terraform | >= 1.1.0 |
| aws | >= 4.0.0 |
| grafana | >= 1.25.0 |
| helm | >= 2.4.1 |
| kubectl | >= 1.14 |
| kubernetes | >= 2.10 |
Providers
| Name | Version |
|---|---|
| aws | >= 4.0.0 |
| grafana | >= 1.25.0 |
| helm | >= 2.4.1 |
Modules
| Name | Source | Version |
|---|---|---|
| fluentbit_logs | ./add-ons/aws-for-fluentbit | n/a |
| helm_addon | github.com/aws-ia/terraform-aws-eks-blueprints//modules/kubernetes-addons/helm-addon | v4.26.0 |
| java_monitoring | ./patterns/java | n/a |
| nginx_monitoring | ./patterns/nginx | n/a |
| operator | ./add-ons/adot-operator | n/a |
Resources
| Name | Type |
|---|---|
| aws_prometheus_rule_group_namespace.alerting_rules | resource |
| aws_prometheus_rule_group_namespace.recording_rules | resource |
| grafana_dashboard.cluster | resource |
| grafana_dashboard.kubelet | resource |
| grafana_dashboard.nodeexp_nodes | resource |
| grafana_dashboard.nodes | resource |
| grafana_dashboard.nsworkload | resource |
| grafana_dashboard.workloads | resource |
| helm_release.kube_state_metrics | resource |
| helm_release.prometheus_node_exporter | resource |
| aws_caller_identity.current | data source |
| aws_eks_cluster.eks_cluster | data source |
| aws_partition.current | data source |
| aws_region.current | data source |
Inputs
| Name | Description | Type | Default | Required |
|---|---|---|---|---|
| custom_metrics_config | Configuration object to enable custom metrics collection | object({ |
{ |
no |
| dashboards_folder_id | Grafana folder ID for automatic dashboards | string |
n/a | yes |
| eks_cluster_id | EKS Cluster Id | string |
n/a | yes |
| enable_alerting_rules | Enables or disables Managed Prometheus alerting rules | bool |
true |
no |
| enable_amazon_eks_adot | Enables the ADOT Operator on the EKS Cluster | bool |
true |
no |
| enable_cert_manager | Allow reusing an existing installation of cert-manager | bool |
true |
no |
| enable_custom_metrics | Allows additional metrics collection for config elements in the custom_metrics_config config object. Automatic dashboards are not included |
bool |
false |
no |
| enable_dashboards | Enables or disables curated dashboards | bool |
true |
no |
| enable_java | Enable Java workloads monitoring, alerting and default dashboards | bool |
false |
no |
| enable_kube_state_metrics | Enables or disables Kube State metrics exporter. Disabling this might affect some data in the dashboards | bool |
true |
no |
| enable_logs | Using AWS For FluentBit to collect cluster and application logs to Amazon CloudWatch | bool |
true |
no |
| enable_nginx | Enable NGINX workloads monitoring, alerting and default dashboards | bool |
false |
no |
| enable_node_exporter | Enables or disables Node exporter. Disabling this might affect some data in the dashboards | bool |
true |
no |
| enable_recording_rules | Enables or disables Managed Prometheus recording rules | bool |
true |
no |
| enable_tracing | (Experimental) Enables tracing with AWS X-Ray. This changes the deploy mode of the collector to daemon set. Requirement: adot add-on <= 0.58-build.0 | bool |
false |
no |
| helm_config | Helm Config for Prometheus | any |
{} |
no |
| irsa_iam_permissions_boundary | IAM permissions boundary for IRSA roles | string |
null |
no |
| irsa_iam_role_path | IAM role path for IRSA roles | string |
"/" |
no |
| java_config | Configuration object for Java/JMX monitoring | object({ |
{ |
no |
| ksm_config | Kube State metrics configuration | object({ |
{ |
no |
| logs_config | Configuration object for logs collection | object({ |
{ |
no |
| managed_prometheus_workspace_endpoint | Amazon Managed Prometheus Workspace Endpoint | string |
"" |
no |
| managed_prometheus_workspace_id | Amazon Managed Prometheus Workspace ID | string |
null |
no |
| managed_prometheus_workspace_region | Amazon Managed Prometheus Workspace's Region | string |
null |
no |
| ne_config | Node exporter configuration | object({ |
{ |
no |
| nginx_config | Configuration object for NGINX monitoring | object({ |
{ |
no |
| prometheus_config | Controls default values such as scrape interval, timeouts and ports globally | object({ |
{ |
no |
| tags | Additional tags (e.g. map('BusinessUnit,XYZ) |
map(string) |
{} |
no |
| tracing_config | Configuration object for traces collection to AWS X-Ray | object({ |
{ |
no |
Outputs
| Name | Description |
|---|---|
| eks_cluster_id | EKS Cluster Id |
| eks_cluster_version | EKS Cluster version |
| grafana_dashboard_urls | URLs for dashboards created |
Troubleshooting
When you upgrade the eks-monitoring module from v2.1.0 or earlier, the following error may occur.
Error: cannot patch "prometheus-node-exporter" with kind DaemonSet: DaemonSet.apps "prometheus-node-exporter" is invalid: spec.selector: Invalid value: v1.LabelSelector{MatchLabels:map[string]string{"app.kubernetes.io/instance":"prometheus-node-exporter", "app.kubernetes.io/name":"prometheus-node-exporter"}, MatchExpressions:[]v1.LabelSelectorRequirement(nil)}: field is immutable
This is due to the upgrade of the node-exporter chart from v2 to v4. Manually delete the node-exporter's DaemonSet as described in the link here, and then apply.
kubectl -n prometheus-node-exporter delete daemonset -l app=prometheus-node-exporter
terraform apply