Files
terraform-aws-observability…/modules/eks-monitoring
Vijay Chintalapati caeebe0e1a Make creation of Grafana Data Source and Folder configurable (#145)
* Made the cluster variable visible in dashboards

* Made the repo reusable with multiple EKS clusters

* Relabled the input variable from create_grafana_data_source to create_prometheus_data_source
2023-04-06 11:04:29 +02:00
..

Infrastructure monitoring

This module provides EKS cluster monitoring with the following resources:

  • AWS Distro For OpenTelemetry Operator and Collector for Metrics and Traces
  • Logs with AWS for FluentBit
  • AWS Managed Grafana Dashboard and data source
  • Alerts and recording rules with AWS Managed Service for Prometheus

This module makes use of the open source kube-prometheus-stack

Requirements

Name Version
terraform >= 1.1.0
aws >= 4.0.0
grafana >= 1.25.0
helm >= 2.4.1
kubectl >= 1.14
kubernetes >= 2.10

Providers

Name Version
aws >= 4.0.0
grafana >= 1.25.0
helm >= 2.4.1

Modules

Name Source Version
fluentbit_logs ./add-ons/aws-for-fluentbit n/a
helm_addon github.com/aws-ia/terraform-aws-eks-blueprints//modules/kubernetes-addons/helm-addon v4.26.0
java_monitoring ./patterns/java n/a
nginx_monitoring ./patterns/nginx n/a
operator ./add-ons/adot-operator n/a

Resources

Name Type
aws_prometheus_rule_group_namespace.alerting_rules resource
aws_prometheus_rule_group_namespace.recording_rules resource
grafana_dashboard.cluster resource
grafana_dashboard.kubelet resource
grafana_dashboard.nodeexp_nodes resource
grafana_dashboard.nodes resource
grafana_dashboard.nsworkload resource
grafana_dashboard.workloads resource
helm_release.kube_state_metrics resource
helm_release.prometheus_node_exporter resource
aws_caller_identity.current data source
aws_eks_cluster.eks_cluster data source
aws_partition.current data source
aws_region.current data source

Inputs

Name Description Type Default Required
custom_metrics_config Configuration object to enable custom metrics collection
object({
ports = list(number)
# paths = optional(list(string), ["/metrics"])
# list of samples to be dropped by label prefix, ex: go_ -> discards go_.*
dropped_series_prefixes = list(string)
})
{
"dropped_series_prefixes": [
"unspecified"
],
"ports": []
}
no
dashboards_folder_id Grafana folder ID for automatic dashboards string n/a yes
eks_cluster_id EKS Cluster Id string n/a yes
enable_alerting_rules Enables or disables Managed Prometheus alerting rules bool true no
enable_amazon_eks_adot Enables the ADOT Operator on the EKS Cluster bool true no
enable_cert_manager Allow reusing an existing installation of cert-manager bool true no
enable_custom_metrics Allows additional metrics collection for config elements in the custom_metrics_config config object. Automatic dashboards are not included bool false no
enable_dashboards Enables or disables curated dashboards bool true no
enable_java Enable Java workloads monitoring, alerting and default dashboards bool false no
enable_kube_state_metrics Enables or disables Kube State metrics exporter. Disabling this might affect some data in the dashboards bool true no
enable_logs Using AWS For FluentBit to collect cluster and application logs to Amazon CloudWatch bool true no
enable_nginx Enable NGINX workloads monitoring, alerting and default dashboards bool false no
enable_node_exporter Enables or disables Node exporter. Disabling this might affect some data in the dashboards bool true no
enable_recording_rules Enables or disables Managed Prometheus recording rules bool true no
enable_tracing (Experimental) Enables tracing with AWS X-Ray. This changes the deploy mode of the collector to daemon set. Requirement: adot add-on <= 0.58-build.0 bool false no
helm_config Helm Config for Prometheus any {} no
irsa_iam_permissions_boundary IAM permissions boundary for IRSA roles string null no
irsa_iam_role_path IAM role path for IRSA roles string "/" no
java_config Configuration object for Java/JMX monitoring
object({
enable_alerting_rules = bool
enable_recording_rules = bool
scrape_sample_limit = number
})
{
"enable_alerting_rules": true,
"enable_recording_rules": true,
"scrape_sample_limit": 1000
}
no
ksm_config Kube State metrics configuration
object({
create_namespace = bool
k8s_namespace = string
helm_chart_name = string
helm_chart_version = string
helm_release_name = string
helm_repo_url = string
helm_settings = map(string)
helm_values = map(any)

scrape_interval = string
scrape_timeout = string
})
{
"create_namespace": true,
"helm_chart_name": "kube-state-metrics",
"helm_chart_version": "4.24.0",
"helm_release_name": "kube-state-metrics",
"helm_repo_url": "https://prometheus-community.github.io/helm-charts",
"helm_settings": {},
"helm_values": {},
"k8s_namespace": "kube-system",
"scrape_interval": "60s",
"scrape_timeout": "15s"
}
no
logs_config Configuration object for logs collection
object({
cw_log_retention_days = number
})
{
"cw_log_retention_days": 90
}
no
managed_prometheus_workspace_endpoint Amazon Managed Prometheus Workspace Endpoint string "" no
managed_prometheus_workspace_id Amazon Managed Prometheus Workspace ID string null no
managed_prometheus_workspace_region Amazon Managed Prometheus Workspace's Region string null no
ne_config Node exporter configuration
object({
create_namespace = bool
k8s_namespace = string
helm_chart_name = string
helm_chart_version = string
helm_release_name = string
helm_repo_url = string
helm_settings = map(string)
helm_values = map(any)

scrape_interval = string
scrape_timeout = string
})
{
"create_namespace": true,
"helm_chart_name": "prometheus-node-exporter",
"helm_chart_version": "4.14.0",
"helm_release_name": "prometheus-node-exporter",
"helm_repo_url": "https://prometheus-community.github.io/helm-charts",
"helm_settings": {},
"helm_values": {},
"k8s_namespace": "prometheus-node-exporter",
"scrape_interval": "60s",
"scrape_timeout": "60s"
}
no
nginx_config Configuration object for NGINX monitoring
object({
enable_alerting_rules = bool
scrape_sample_limit = number
prometheus_metrics_endpoint = string
})
{
"enable_alerting_rules": true,
"prometheus_metrics_endpoint": "metrics",
"scrape_sample_limit": 1000
}
no
prometheus_config Controls default values such as scrape interval, timeouts and ports globally
object({
global_scrape_interval = string
global_scrape_timeout = string
})
{
"global_scrape_interval": "60s",
"global_scrape_timeout": "15s"
}
no
tags Additional tags (e.g. map('BusinessUnit,XYZ) map(string) {} no
tracing_config Configuration object for traces collection to AWS X-Ray
object({
otlp_grpc_endpoint = string
otlp_http_endpoint = string
send_batch_size = number
timeout = string
})
{
"otlp_grpc_endpoint": "0.0.0.0:4317",
"otlp_http_endpoint": "0.0.0.0:4318",
"send_batch_size": 50,
"timeout": "30s"
}
no

Outputs

Name Description
eks_cluster_id EKS Cluster Id
eks_cluster_version EKS Cluster version
grafana_dashboard_urls URLs for dashboards created

Troubleshooting

When you upgrade the eks-monitoring module from v2.1.0 or earlier, the following error may occur.

Error: cannot patch "prometheus-node-exporter" with kind DaemonSet: DaemonSet.apps "prometheus-node-exporter" is invalid: spec.selector: Invalid value: v1.LabelSelector{MatchLabels:map[string]string{"app.kubernetes.io/instance":"prometheus-node-exporter", "app.kubernetes.io/name":"prometheus-node-exporter"}, MatchExpressions:[]v1.LabelSelectorRequirement(nil)}: field is immutable

This is due to the upgrade of the node-exporter chart from v2 to v4. Manually delete the node-exporter's DaemonSet as described in the link here, and then apply.

kubectl -n prometheus-node-exporter delete daemonset -l app=prometheus-node-exporter
terraform apply