Compose EKS monitoring modules (#115)

* Move modules around

* Update amp billing source

* Merge Java monitoring to EKS

* Update docs

* Merge nginx pattern

* Pre-commit

* Add save and test URL output

* Move EKS dependencies to EKS monitoring module

* update docs

* Update examples and docs

* Add java doc

* Add NGINX doc

* Update nginx doc

* Fix amp monitoring example path

* Fix pre-commit

* Todo: move to main after merge

* Update docs, fix tags
This commit is contained in:
Rodrigue Koffi
2023-02-20 18:36:08 +01:00
committed by GitHub
parent fe83579997
commit daed34db80
99 changed files with 887 additions and 1477 deletions
@@ -33,6 +33,9 @@ This module is inspired from the open source [kube-prometheus-stack](https://git
| Name | Source | Version |
|------|--------|---------|
| <a name="module_helm_addon"></a> [helm\_addon](#module\_helm\_addon) | github.com/aws-ia/terraform-aws-eks-blueprints//modules/kubernetes-addons/helm-addon | v4.13.1 |
| <a name="module_java_monitoring"></a> [java\_monitoring](#module\_java\_monitoring) | ./patterns/java | n/a |
| <a name="module_nginx_monitoring"></a> [nginx\_monitoring](#module\_nginx\_monitoring) | ./patterns/nginx | n/a |
| <a name="module_operator"></a> [operator](#module\_operator) | ./add-ons/adot-operator | n/a |
## Resources
@@ -61,20 +64,25 @@ This module is inspired from the open source [kube-prometheus-stack](https://git
| <a name="input_dashboards_folder_id"></a> [dashboards\_folder\_id](#input\_dashboards\_folder\_id) | Grafana folder ID for automatic dashboards | `string` | n/a | yes |
| <a name="input_eks_cluster_id"></a> [eks\_cluster\_id](#input\_eks\_cluster\_id) | EKS Cluster Id | `string` | n/a | yes |
| <a name="input_enable_alerting_rules"></a> [enable\_alerting\_rules](#input\_enable\_alerting\_rules) | Enables or disables Managed Prometheus alerting rules | `bool` | `true` | no |
| <a name="input_enable_amazon_eks_adot"></a> [enable\_amazon\_eks\_adot](#input\_enable\_amazon\_eks\_adot) | Enables the ADOT Operator on the EKS Cluster | `bool` | `true` | no |
| <a name="input_enable_cert_manager"></a> [enable\_cert\_manager](#input\_enable\_cert\_manager) | Allow reusing an existing installation of cert-manager | `bool` | `true` | no |
| <a name="input_enable_custom_metrics"></a> [enable\_custom\_metrics](#input\_enable\_custom\_metrics) | Allows additional metrics collection for config elements in the `custom_metrics_config` config object. Automatic dashboards are not included | `bool` | `false` | no |
| <a name="input_enable_dashboards"></a> [enable\_dashboards](#input\_enable\_dashboards) | Enables or disables curated dashboards | `bool` | `true` | no |
| <a name="input_enable_java"></a> [enable\_java](#input\_enable\_java) | Enable Java workloads monitoring, alerting and default dashboards | `bool` | `false` | no |
| <a name="input_enable_kube_state_metrics"></a> [enable\_kube\_state\_metrics](#input\_enable\_kube\_state\_metrics) | Enables or disables Kube State metrics exporter. Disabling this might affect some data in the dashboards | `bool` | `true` | no |
| <a name="input_enable_nginx"></a> [enable\_nginx](#input\_enable\_nginx) | Enable NGINX workloads monitoring, alerting and default dashboards | `bool` | `false` | no |
| <a name="input_enable_node_exporter"></a> [enable\_node\_exporter](#input\_enable\_node\_exporter) | Enables or disables Node exporter. Disabling this might affect some data in the dashboards | `bool` | `true` | no |
| <a name="input_enable_recording_rules"></a> [enable\_recording\_rules](#input\_enable\_recording\_rules) | Enables or disables Managed Prometheus recording rules. Disabling this might affect some data in the dashboards | `bool` | `true` | no |
| <a name="input_enable_tracing"></a> [enable\_tracing](#input\_enable\_tracing) | (Experimental) Enables tracing with AWS X-Ray. This changes the deploy mode of the collector to daemon set. Requirement: adot add-on <= 0.58-build.0 | `bool` | `false` | no |
| <a name="input_helm_config"></a> [helm\_config](#input\_helm\_config) | Helm Config for Prometheus | `any` | `{}` | no |
| <a name="input_irsa_iam_permissions_boundary"></a> [irsa\_iam\_permissions\_boundary](#input\_irsa\_iam\_permissions\_boundary) | IAM permissions boundary for IRSA roles | `string` | `null` | no |
| <a name="input_irsa_iam_role_path"></a> [irsa\_iam\_role\_path](#input\_irsa\_iam\_role\_path) | IAM role path for IRSA roles | `string` | `"/"` | no |
| <a name="input_java_config"></a> [java\_config](#input\_java\_config) | Configuration object for Java/JMX monitoring | <pre>object({<br> enable_alerting_rules = bool<br> scrape_sample_limit = number<br> })</pre> | <pre>{<br> "enable_alerting_rules": true,<br> "scrape_sample_limit": 1000<br>}</pre> | no |
| <a name="input_ksm_config"></a> [ksm\_config](#input\_ksm\_config) | Kube State metrics configuration | <pre>object({<br> create_namespace = bool<br> k8s_namespace = string<br> helm_chart_name = string<br> helm_chart_version = string<br> helm_release_name = string<br> helm_repo_url = string<br> helm_settings = map(string)<br> helm_values = map(any)<br><br> scrape_interval = string<br> scrape_timeout = string<br> })</pre> | <pre>{<br> "create_namespace": true,<br> "helm_chart_name": "kube-state-metrics",<br> "helm_chart_version": "4.24.0",<br> "helm_release_name": "kube-state-metrics",<br> "helm_repo_url": "https://prometheus-community.github.io/helm-charts",<br> "helm_settings": {},<br> "helm_values": {},<br> "k8s_namespace": "kube-system",<br> "scrape_interval": "60s",<br> "scrape_timeout": "15s"<br>}</pre> | no |
| <a name="input_managed_prometheus_workspace_endpoint"></a> [managed\_prometheus\_workspace\_endpoint](#input\_managed\_prometheus\_workspace\_endpoint) | Amazon Managed Prometheus Workspace Endpoint | `string` | `""` | no |
| <a name="input_managed_prometheus_workspace_id"></a> [managed\_prometheus\_workspace\_id](#input\_managed\_prometheus\_workspace\_id) | Amazon Managed Prometheus Workspace ID | `string` | `null` | no |
| <a name="input_managed_prometheus_workspace_region"></a> [managed\_prometheus\_workspace\_region](#input\_managed\_prometheus\_workspace\_region) | Amazon Managed Prometheus Workspace's Region | `string` | `null` | no |
| <a name="input_ne_config"></a> [ne\_config](#input\_ne\_config) | Node exporter configuration | <pre>object({<br> create_namespace = bool<br> k8s_namespace = string<br> helm_chart_name = string<br> helm_chart_version = string<br> helm_release_name = string<br> helm_repo_url = string<br> helm_settings = map(string)<br> helm_values = map(any)<br><br> scrape_interval = string<br> scrape_timeout = string<br> })</pre> | <pre>{<br> "create_namespace": true,<br> "helm_chart_name": "prometheus-node-exporter",<br> "helm_chart_version": "2.0.3",<br> "helm_release_name": "prometheus-node-exporter",<br> "helm_repo_url": "https://prometheus-community.github.io/helm-charts",<br> "helm_settings": {},<br> "helm_values": {},<br> "k8s_namespace": "prometheus-node-exporter",<br> "scrape_interval": "60s",<br> "scrape_timeout": "60s"<br>}</pre> | no |
| <a name="input_nginx_config"></a> [nginx\_config](#input\_nginx\_config) | Configuration object for NGINX monitoring | <pre>object({<br> enable_alerting_rules = bool<br> scrape_sample_limit = number<br> prometheus_metrics_endpoint = string<br> })</pre> | <pre>{<br> "enable_alerting_rules": true,<br> "prometheus_metrics_endpoint": "metrics",<br> "scrape_sample_limit": 1000<br>}</pre> | no |
| <a name="input_prometheus_config"></a> [prometheus\_config](#input\_prometheus\_config) | Controls default values such as scrape interval, timeouts and ports globally | <pre>object({<br> global_scrape_interval = string<br> global_scrape_timeout = string<br> })</pre> | <pre>{<br> "global_scrape_interval": "60s",<br> "global_scrape_timeout": "15s"<br>}</pre> | no |
| <a name="input_tags"></a> [tags](#input\_tags) | Additional tags (e.g. `map('BusinessUnit`,`XYZ`) | `map(string)` | `{}` | no |
| <a name="input_tracing_config"></a> [tracing\_config](#input\_tracing\_config) | Configuration object for traces collection to AWS X-Ray | <pre>object({<br> otlp_grpc_endpoint = string<br> otlp_http_endpoint = string<br> send_batch_size = number<br> timeout = string<br> })</pre> | <pre>{<br> "otlp_grpc_endpoint": "0.0.0.0:4317",<br> "otlp_http_endpoint": "0.0.0.0:4318",<br> "send_batch_size": 50,<br> "timeout": "30s"<br>}</pre> | no |
@@ -83,5 +91,7 @@ This module is inspired from the open source [kube-prometheus-stack](https://git
| Name | Description |
|------|-------------|
| <a name="output_eks_cluster_id"></a> [eks\_cluster\_id](#output\_eks\_cluster\_id) | EKS Cluster Id |
| <a name="output_eks_cluster_version"></a> [eks\_cluster\_version](#output\_eks\_cluster\_version) | EKS Cluster version |
| <a name="output_grafana_dashboard_urls"></a> [grafana\_dashboard\_urls](#output\_grafana\_dashboard\_urls) | URLs for dashboards created |
<!-- END OF PRE-COMMIT-TERRAFORM DOCS HOOK -->
@@ -14,6 +14,7 @@ locals {
eks_oidc_issuer_url = replace(data.aws_eks_cluster.eks_cluster.identity[0].oidc[0].issuer, "https://", "")
eks_cluster_endpoint = data.aws_eks_cluster.eks_cluster.endpoint
eks_cluster_version = data.aws_eks_cluster.eks_cluster.version
context = {
aws_caller_identity_account_id = data.aws_caller_identity.current.account_id
@@ -1,3 +1,12 @@
module "operator" {
source = "./add-ons/adot-operator"
count = var.enable_amazon_eks_adot ? 1 : 0
enable_cert_manager = var.enable_cert_manager
kubernetes_version = local.eks_cluster_version
addon_context = local.context
}
resource "helm_release" "kube_state_metrics" {
count = var.enable_kube_state_metrics ? 1 : 0
chart = var.ksm_config.helm_chart_name
@@ -104,6 +113,26 @@ module "helm_addon" {
{
name = "customMetricsDroppedSeriesPrefixes"
value = format("(%s.*)$", join(".*|", var.custom_metrics_config.dropped_series_prefixes))
},
{
name = "enable_java"
value = var.enable_java
},
{
name = "javaScrapeSampleLimit"
value = var.java_config.scrape_sample_limit
},
{
name = "enable_nginx"
value = var.enable_nginx
},
{
name = "nginxScrapeSampleLimit"
value = var.nginx_config.scrape_sample_limit
},
{
name = "nginxPrometheusMetricsEndpoint"
value = var.nginx_config.prometheus_metrics_endpoint
}
]
@@ -119,4 +148,24 @@ module "helm_addon" {
}
addon_context = local.context
depends_on = [module.operator]
}
module "java_monitoring" {
source = "./patterns/java"
count = var.enable_java ? 1 : 0
managed_prometheus_workspace_id = var.managed_prometheus_workspace_id
enable_alerting_rules = var.java_config.enable_alerting_rules
dashboards_folder_id = var.dashboards_folder_id
}
module "nginx_monitoring" {
source = "./patterns/nginx"
count = var.enable_nginx ? 1 : 0
managed_prometheus_workspace_id = var.managed_prometheus_workspace_id
enable_alerting_rules = var.nginx_config.enable_alerting_rules
dashboards_folder_id = var.dashboards_folder_id
}
@@ -1703,6 +1703,72 @@ spec:
- source_labels: [ __name__ ]
regex: '{{ .Values.customMetricsDroppedSeriesPrefixes }}'
action: drop
{{ if .Values.enableJava }}
- job_name: 'kubernetes-java-jmx'
sample_limit: {{ .Values.javaScrapeSampleLimit }}
kubernetes_sd_configs:
- role: pod
relabel_configs:
- source_labels: [ __address__ ]
action: keep
regex: '.*:9404$'
- action: labelmap
regex: __meta_kubernetes_pod_label_(.+)
- action: replace
source_labels: [ __meta_kubernetes_namespace ]
target_label: Namespace
- source_labels: [ __meta_kubernetes_pod_name ]
action: replace
target_label: pod_name
- action: replace
source_labels: [ __meta_kubernetes_pod_container_name ]
target_label: container_name
- action: replace
source_labels: [ __meta_kubernetes_pod_controller_kind ]
target_label: pod_controller_kind
- action: replace
source_labels: [ __meta_kubernetes_pod_phase ]
target_label: pod_controller_phase
metric_relabel_configs:
- source_labels: [ __name__ ]
regex: 'jvm_gc_collection_seconds.*'
action: drop
{{ end }}
{{ if .Values.enableNginx }}
- job_name: 'kubernetes-nginx'
sample_limit: {{ .Values.nginxScrapeSampleLimit }}
metrics_path: /{{ .Values.nginxPrometheusMetricsEndpoint }}
kubernetes_sd_configs:
- role: pod
relabel_configs:
- source_labels: [ __address__ ]
action: keep
regex: '.*:10254$'
- source_labels: [__meta_kubernetes_pod_container_name]
target_label: container
action: replace
- source_labels: [__meta_kubernetes_pod_node_name]
target_label: host
action: replace
- source_labels: [__meta_kubernetes_namespace]
target_label: namespace
action: replace
metric_relabel_configs:
- source_labels: [__name__]
regex: 'go_memstats.*'
action: drop
- source_labels: [__name__]
regex: 'go_gc.*'
action: drop
- source_labels: [__name__]
regex: 'go_threads'
action: drop
- regex: exported_host
action: labeldrop
{{ end }}
exporters:
{{ if .Values.enableTracing }}
awsxray:
@@ -14,3 +14,10 @@ tracingSendBatchSize: ${tracing_send_batch_size}
enableCustomMetrics: ${enable_custom_metrics}
customMetricsPorts: ${custom_metrics_ports}
customMetricsDroppedSeriesPrefixes: ${custom_metrics_dropped_series_prefixes}
enableJava: ${enable_java}
javaScrapeSampleLimit: ${java_scrape_sample_limit}
enableNginx: ${enable_nginx}
nginxScrapeSampleLimit: ${nginx_scrape_sample_limit}
nginxPrometheusMetricsEndpoint: ${nginx_prometheus_metrics_endpoint}
+21
View File
@@ -0,0 +1,21 @@
output "grafana_dashboard_urls" {
value = [concat(
grafana_dashboard.workloads[*].url,
grafana_dashboard.nodes[*].url,
grafana_dashboard.nsworkload[*].url,
grafana_dashboard.kubelet[*].url,
grafana_dashboard.cluster[*].url,
flatten(module.java_monitoring[*].grafana_dashboard_urls),
flatten(module.nginx_monitoring[*].grafana_dashboard_urls),
)]
description = "URLs for dashboards created"
}
output "eks_cluster_version" {
description = "EKS Cluster version"
value = data.aws_eks_cluster.eks_cluster.version
}
output "eks_cluster_id" {
description = "EKS Cluster Id"
value = var.eks_cluster_id
}
@@ -0,0 +1,52 @@
# Java patterns module
Provides monitoring for Java based workloads with the following resources:
- AWS Managed Grafana Dashboard and data source
- Alerts and recording rules with AWS Managed Service for Prometheus
<!-- BEGINNING OF PRE-COMMIT-TERRAFORM DOCS HOOK -->
## Requirements
| Name | Version |
|------|---------|
| <a name="requirement_terraform"></a> [terraform](#requirement\_terraform) | >= 1.1.0 |
| <a name="requirement_aws"></a> [aws](#requirement\_aws) | >= 4.0.0 |
| <a name="requirement_grafana"></a> [grafana](#requirement\_grafana) | >= 1.25.0 |
| <a name="requirement_helm"></a> [helm](#requirement\_helm) | >= 2.4.1 |
| <a name="requirement_kubectl"></a> [kubectl](#requirement\_kubectl) | >= 1.14 |
| <a name="requirement_kubernetes"></a> [kubernetes](#requirement\_kubernetes) | >= 2.10 |
## Providers
| Name | Version |
|------|---------|
| <a name="provider_aws"></a> [aws](#provider\_aws) | >= 4.0.0 |
| <a name="provider_grafana"></a> [grafana](#provider\_grafana) | >= 1.25.0 |
## Modules
No modules.
## Resources
| Name | Type |
|------|------|
| [aws_prometheus_rule_group_namespace.alerting_rules](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/resources/prometheus_rule_group_namespace) | resource |
| [aws_prometheus_rule_group_namespace.recording_rules](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/resources/prometheus_rule_group_namespace) | resource |
| [grafana_dashboard.this](https://registry.terraform.io/providers/grafana/grafana/latest/docs/resources/dashboard) | resource |
## Inputs
| Name | Description | Type | Default | Required |
|------|-------------|------|---------|:--------:|
| <a name="input_dashboards_folder_id"></a> [dashboards\_folder\_id](#input\_dashboards\_folder\_id) | Grafana folder ID for automatic dashboards | `string` | n/a | yes |
| <a name="input_enable_alerting_rules"></a> [enable\_alerting\_rules](#input\_enable\_alerting\_rules) | Enables or disables Managed Prometheus alerting rules | `bool` | `true` | no |
| <a name="input_managed_prometheus_workspace_id"></a> [managed\_prometheus\_workspace\_id](#input\_managed\_prometheus\_workspace\_id) | Amazon Managed Prometheus Workspace ID | `string` | `null` | no |
## Outputs
| Name | Description |
|------|-------------|
| <a name="output_grafana_dashboard_urls"></a> [grafana\_dashboard\_urls](#output\_grafana\_dashboard\_urls) | URLs for dashboards created |
<!-- END OF PRE-COMMIT-TERRAFORM DOCS HOOK -->
@@ -1837,4 +1837,4 @@
"uid": "m9mHfAy7ks",
"version": 4,
"weekStart": ""
}
}
@@ -0,0 +1,36 @@
resource "aws_prometheus_rule_group_namespace" "recording_rules" {
name = "accelerator-java-rules"
workspace_id = var.managed_prometheus_workspace_id
data = <<EOF
groups:
- name: default-metric
rules:
- record: metric:recording_rule
expr: avg(rate(container_cpu_usage_seconds_total[5m]))
EOF
}
resource "aws_prometheus_rule_group_namespace" "alerting_rules" {
count = var.enable_alerting_rules ? 1 : 0
name = "accelerator-java-alerting"
workspace_id = var.managed_prometheus_workspace_id
data = <<EOF
groups:
- name: default-alert
rules:
- alert: metric:alerting_rule
expr: jvm_memory_bytes_used{job="java", area="heap"} / jvm_memory_bytes_max * 100 > 80
for: 1m
labels:
severity: warning
annotations:
summary: "JVM heap warning"
description: "JVM heap of instance `{{$labels.instance}}` from application `{{$labels.application}}` is above 80% for one minute. (current=`{{$value}}%`)"
EOF
}
resource "grafana_dashboard" "this" {
folder = var.dashboards_folder_id
config_json = file("${path.module}/dashboards/default.json")
}
@@ -0,0 +1,16 @@
variable "enable_alerting_rules" {
description = "Enables or disables Managed Prometheus alerting rules"
type = bool
default = true
}
variable "managed_prometheus_workspace_id" {
description = "Amazon Managed Prometheus Workspace ID"
type = string
default = null
}
variable "dashboards_folder_id" {
description = "Grafana folder ID for automatic dashboards"
type = string
}
@@ -26,9 +26,7 @@ It provides the following resources:
## Modules
| Name | Source | Version |
|------|--------|---------|
| <a name="module_helm_addon"></a> [helm\_addon](#module\_helm\_addon) | github.com/aws-ia/terraform-aws-eks-blueprints//modules/kubernetes-addons/helm-addon | v4.13.1 |
No modules.
## Resources
@@ -36,27 +34,14 @@ It provides the following resources:
|------|------|
| [aws_prometheus_rule_group_namespace.alerting_rules](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/resources/prometheus_rule_group_namespace) | resource |
| [grafana_dashboard.workloads](https://registry.terraform.io/providers/grafana/grafana/latest/docs/resources/dashboard) | resource |
| [aws_caller_identity.current](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/data-sources/caller_identity) | data source |
| [aws_eks_cluster.eks_cluster](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/data-sources/eks_cluster) | data source |
| [aws_partition.current](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/data-sources/partition) | data source |
| [aws_region.current](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/data-sources/region) | data source |
## Inputs
| Name | Description | Type | Default | Required |
|------|-------------|------|---------|:--------:|
| <a name="input_config"></a> [config](#input\_config) | Helm Config for Prometheus | `any` | `{}` | no |
| <a name="input_dashboards_folder_id"></a> [dashboards\_folder\_id](#input\_dashboards\_folder\_id) | Grafana folder ID for automatic dashboards | `string` | n/a | yes |
| <a name="input_eks_cluster_id"></a> [eks\_cluster\_id](#input\_eks\_cluster\_id) | EKS Cluster Id | `string` | n/a | yes |
| <a name="input_enable_alerting_rules"></a> [enable\_alerting\_rules](#input\_enable\_alerting\_rules) | Enables or disables Managed Prometheus alerting rules | `bool` | `true` | no |
| <a name="input_enable_dashboards"></a> [enable\_dashboards](#input\_enable\_dashboards) | Enables or disables curated dashboards | `bool` | `true` | no |
| <a name="input_helm_config"></a> [helm\_config](#input\_helm\_config) | Helm Config for Prometheus | `any` | `{}` | no |
| <a name="input_irsa_iam_permissions_boundary"></a> [irsa\_iam\_permissions\_boundary](#input\_irsa\_iam\_permissions\_boundary) | IAM permissions boundary for IRSA roles | `string` | `null` | no |
| <a name="input_irsa_iam_role_path"></a> [irsa\_iam\_role\_path](#input\_irsa\_iam\_role\_path) | IAM role path for IRSA roles | `string` | `"/"` | no |
| <a name="input_managed_prometheus_workspace_endpoint"></a> [managed\_prometheus\_workspace\_endpoint](#input\_managed\_prometheus\_workspace\_endpoint) | Amazon Managed Prometheus Workspace Endpoint | `string` | `""` | no |
| <a name="input_managed_prometheus_workspace_id"></a> [managed\_prometheus\_workspace\_id](#input\_managed\_prometheus\_workspace\_id) | Amazon Managed Prometheus Workspace ID | `string` | `null` | no |
| <a name="input_managed_prometheus_workspace_region"></a> [managed\_prometheus\_workspace\_region](#input\_managed\_prometheus\_workspace\_region) | Amazon Managed Prometheus Workspace's Region | `string` | `null` | no |
| <a name="input_tags"></a> [tags](#input\_tags) | Additional tags (e.g. `map('BusinessUnit`,`XYZ`) | `map(string)` | `{}` | no |
## Outputs
@@ -1,7 +1,3 @@
################################################################################################################################################
# Alerting rules ###############################################################################################################################
################################################################################################################################################
resource "aws_prometheus_rule_group_namespace" "alerting_rules" {
count = var.enable_alerting_rules ? 1 : 0
@@ -41,3 +37,8 @@ groups:
description: "Nginx p99 latency is higher than 3 seconds\n VALUE = {{ $value }}\n LABELS = {{ $labels }}"
EOF
}
resource "grafana_dashboard" "workloads" {
folder = var.dashboards_folder_id
config_json = file("${path.module}/dashboards/nginx.json")
}
@@ -0,0 +1,16 @@
variable "managed_prometheus_workspace_id" {
description = "Amazon Managed Prometheus Workspace ID"
type = string
default = null
}
variable "dashboards_folder_id" {
type = string
description = "Grafana folder ID for automatic dashboards"
}
variable "enable_alerting_rules" {
type = bool
default = true
description = "Enables or disables Managed Prometheus alerting rules"
}
@@ -5,7 +5,6 @@
################################################################################################################################################
resource "aws_prometheus_rule_group_namespace" "recording_rules" {
count = var.enable_recording_rules ? 1 : 0
name = "accelerator-infra-rules"
workspace_id = var.managed_prometheus_workspace_id
data = <<EOF
@@ -3,6 +3,18 @@ variable "eks_cluster_id" {
type = string
}
variable "enable_amazon_eks_adot" {
description = "Enables the ADOT Operator on the EKS Cluster"
type = bool
default = true
}
variable "enable_cert_manager" {
description = "Allow reusing an existing installation of cert-manager"
type = bool
default = true
}
variable "helm_config" {
description = "Helm Config for Prometheus"
type = any
@@ -44,12 +56,6 @@ variable "dashboards_folder_id" {
type = string
}
variable "enable_recording_rules" {
description = "Enables or disables Managed Prometheus recording rules. Disabling this might affect some data in the dashboards"
type = bool
default = true
}
variable "enable_alerting_rules" {
description = "Enables or disables Managed Prometheus alerting rules"
type = bool
@@ -201,3 +207,43 @@ variable "custom_metrics_config" {
dropped_series_prefixes = ["unspecified"]
}
}
variable "enable_java" {
description = "Enable Java workloads monitoring, alerting and default dashboards"
type = bool
default = false
}
variable "java_config" {
description = "Configuration object for Java/JMX monitoring"
type = object({
enable_alerting_rules = bool
scrape_sample_limit = number
})
default = {
enable_alerting_rules = true
scrape_sample_limit = 1000
}
}
variable "enable_nginx" {
description = "Enable NGINX workloads monitoring, alerting and default dashboards"
type = bool
default = false
}
variable "nginx_config" {
description = "Configuration object for NGINX monitoring"
type = object({
enable_alerting_rules = bool
scrape_sample_limit = number
prometheus_metrics_endpoint = string
})
default = {
enable_alerting_rules = true
scrape_sample_limit = 1000
prometheus_metrics_endpoint = "metrics"
}
}
@@ -9,9 +9,11 @@ locals {
}
resource "grafana_data_source" "cloudwatch" {
type = "cloudwatch"
name = local.name
is_default = true
type = "cloudwatch"
name = local.name
# Giving priority to Managed Prometheus datasources
is_default = false
json_data {
default_region = var.aws_region
sigv4_auth = true
@@ -26,7 +28,7 @@ resource "grafana_dashboard" "this" {
}
module "billing" {
source = "../../workloads/managed-prometheus-monitoring/billing"
source = "./billing"
providers = {
aws = aws.billing_region
}
-10
View File
@@ -1,10 +0,0 @@
output "grafana_dashboard_urls" {
value = [concat(
grafana_dashboard.workloads[*].url,
grafana_dashboard.nodes[*].url,
grafana_dashboard.nsworkload[*].url,
grafana_dashboard.kubelet[*].url,
grafana_dashboard.cluster[*].url,
)]
description = "URLs for dashboards created"
}
-68
View File
@@ -1,68 +0,0 @@
# Java based workloads monitoring
This module provides monitoring for Java based workloads with the following resources:
- AWS Distro For OpenTelemetry Operator and Collector
- AWS Managed Grafana Dashboard and data source
- Alerts and recording rules with AWS Managed Service for Prometheus
<!-- BEGINNING OF PRE-COMMIT-TERRAFORM DOCS HOOK -->
## Requirements
| Name | Version |
|------|---------|
| <a name="requirement_terraform"></a> [terraform](#requirement\_terraform) | >= 1.1.0 |
| <a name="requirement_aws"></a> [aws](#requirement\_aws) | >= 4.0.0 |
| <a name="requirement_grafana"></a> [grafana](#requirement\_grafana) | >= 1.25.0 |
| <a name="requirement_helm"></a> [helm](#requirement\_helm) | >= 2.4.1 |
| <a name="requirement_kubectl"></a> [kubectl](#requirement\_kubectl) | >= 1.14 |
| <a name="requirement_kubernetes"></a> [kubernetes](#requirement\_kubernetes) | >= 2.10 |
## Providers
| Name | Version |
|------|---------|
| <a name="provider_aws"></a> [aws](#provider\_aws) | >= 4.0.0 |
| <a name="provider_grafana"></a> [grafana](#provider\_grafana) | >= 1.25.0 |
## Modules
| Name | Source | Version |
|------|--------|---------|
| <a name="module_helm_addon"></a> [helm\_addon](#module\_helm\_addon) | github.com/aws-ia/terraform-aws-eks-blueprints//modules/kubernetes-addons/helm-addon | v4.13.1 |
## Resources
| Name | Type |
|------|------|
| [aws_prometheus_rule_group_namespace.alerting_rules](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/resources/prometheus_rule_group_namespace) | resource |
| [aws_prometheus_rule_group_namespace.recording_rules](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/resources/prometheus_rule_group_namespace) | resource |
| [grafana_dashboard.this](https://registry.terraform.io/providers/grafana/grafana/latest/docs/resources/dashboard) | resource |
| [aws_caller_identity.current](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/data-sources/caller_identity) | data source |
| [aws_eks_cluster.eks_cluster](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/data-sources/eks_cluster) | data source |
| [aws_partition.current](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/data-sources/partition) | data source |
| [aws_region.current](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/data-sources/region) | data source |
## Inputs
| Name | Description | Type | Default | Required |
|------|-------------|------|---------|:--------:|
| <a name="input_dashboards_folder_id"></a> [dashboards\_folder\_id](#input\_dashboards\_folder\_id) | Grafana folder ID for automatic dashboards | `string` | n/a | yes |
| <a name="input_eks_cluster_id"></a> [eks\_cluster\_id](#input\_eks\_cluster\_id) | EKS Cluster Id | `string` | n/a | yes |
| <a name="input_enable_alerting_rules"></a> [enable\_alerting\_rules](#input\_enable\_alerting\_rules) | Enables or disables Managed Prometheus alerting rules | `bool` | `true` | no |
| <a name="input_enable_recording_rules"></a> [enable\_recording\_rules](#input\_enable\_recording\_rules) | Enables or disables Managed Prometheus recording rules. Disabling this might affect some data in the dashboards | `bool` | `true` | no |
| <a name="input_helm_config"></a> [helm\_config](#input\_helm\_config) | Helm Config for Prometheus | `any` | `{}` | no |
| <a name="input_irsa_iam_permissions_boundary"></a> [irsa\_iam\_permissions\_boundary](#input\_irsa\_iam\_permissions\_boundary) | IAM permissions boundary for IRSA roles | `string` | `null` | no |
| <a name="input_irsa_iam_role_path"></a> [irsa\_iam\_role\_path](#input\_irsa\_iam\_role\_path) | IAM role path for IRSA roles | `string` | `"/"` | no |
| <a name="input_managed_prometheus_workspace_endpoint"></a> [managed\_prometheus\_workspace\_endpoint](#input\_managed\_prometheus\_workspace\_endpoint) | Amazon Managed Prometheus Workspace Endpoint | `string` | `""` | no |
| <a name="input_managed_prometheus_workspace_id"></a> [managed\_prometheus\_workspace\_id](#input\_managed\_prometheus\_workspace\_id) | Amazon Managed Prometheus Workspace ID | `string` | `null` | no |
| <a name="input_managed_prometheus_workspace_region"></a> [managed\_prometheus\_workspace\_region](#input\_managed\_prometheus\_workspace\_region) | Amazon Managed Prometheus Workspace's Region | `string` | `null` | no |
| <a name="input_prometheus_config"></a> [prometheus\_config](#input\_prometheus\_config) | Controls default values such as scrape interval, timeouts and ports globally | <pre>object({<br> global_scrape_interval = string<br> global_scrape_timeout = string<br> scrape_sample_limit = number<br> })</pre> | <pre>{<br> "global_scrape_interval": "60s",<br> "global_scrape_timeout": "15s",<br> "scrape_sample_limit": 1000<br>}</pre> | no |
| <a name="input_tags"></a> [tags](#input\_tags) | Additional tags (e.g. `map('BusinessUnit`,`XYZ`) | `map(string)` | `{}` | no |
## Outputs
| Name | Description |
|------|-------------|
| <a name="output_grafana_dashboard_urls"></a> [grafana\_dashboard\_urls](#output\_grafana\_dashboard\_urls) | URLs for dashboards created |
<!-- END OF PRE-COMMIT-TERRAFORM DOCS HOOK -->
-31
View File
@@ -1,31 +0,0 @@
data "aws_partition" "current" {}
data "aws_caller_identity" "current" {}
data "aws_region" "current" {}
data "aws_eks_cluster" "eks_cluster" {
name = var.eks_cluster_id
}
locals {
name = "adot-collector-java"
namespace = try(var.helm_config.namespace, local.name)
eks_oidc_issuer_url = replace(data.aws_eks_cluster.eks_cluster.identity[0].oidc[0].issuer, "https://", "")
eks_cluster_endpoint = data.aws_eks_cluster.eks_cluster.endpoint
context = {
aws_caller_identity_account_id = data.aws_caller_identity.current.account_id
aws_caller_identity_arn = data.aws_caller_identity.current.arn
aws_eks_cluster_endpoint = local.eks_cluster_endpoint
aws_partition_id = data.aws_partition.current.partition
aws_region_name = data.aws_region.current.name
eks_cluster_id = var.eks_cluster_id
eks_oidc_issuer_url = local.eks_oidc_issuer_url
eks_oidc_provider_arn = "arn:${data.aws_partition.current.partition}:iam::${data.aws_caller_identity.current.account_id}:oidc-provider/${local.eks_oidc_issuer_url}"
tags = var.tags
irsa_iam_role_path = var.irsa_iam_role_path
irsa_iam_permissions_boundary = var.irsa_iam_permissions_boundary
}
}
-95
View File
@@ -1,95 +0,0 @@
# deploys collector
module "helm_addon" {
source = "github.com/aws-ia/terraform-aws-eks-blueprints//modules/kubernetes-addons/helm-addon?ref=v4.13.1"
helm_config = merge(
{
name = local.name
chart = "${path.module}/otel-config"
version = "0.2.0"
namespace = local.namespace
description = "ADOT helm Chart deployment configuration"
},
var.helm_config
)
set_values = [
{
name = "ampurl"
value = "${var.managed_prometheus_workspace_endpoint}api/v1/remote_write"
},
{
name = "region"
value = var.managed_prometheus_workspace_region
},
{
name = "ekscluster"
value = local.context.eks_cluster_id
},
{
name = "accountId"
value = local.context.aws_caller_identity_account_id
},
{
name = "globalScrapeInterval"
value = var.prometheus_config.global_scrape_interval
},
{
name = "globalScrapeTimeout"
value = var.prometheus_config.global_scrape_timeout
},
{
name = "scrapeSampleLimit"
value = var.prometheus_config.scrape_sample_limit
}
]
irsa_config = {
create_kubernetes_namespace = try(var.helm_config["create_namespace"], true)
kubernetes_namespace = local.namespace
create_kubernetes_service_account = true
kubernetes_service_account = try(var.helm_config.service_account, local.name)
irsa_iam_policies = ["arn:${data.aws_partition.current.partition}:iam::aws:policy/AmazonPrometheusRemoteWriteAccess"]
}
addon_context = local.context
}
resource "aws_prometheus_rule_group_namespace" "recording_rules" {
count = var.enable_recording_rules ? 1 : 0
name = "accelerator-java-rules"
workspace_id = var.managed_prometheus_workspace_id
data = <<EOF
groups:
- name: default-metric
rules:
- record: metric:recording_rule
expr: avg(rate(container_cpu_usage_seconds_total[5m]))
EOF
}
resource "aws_prometheus_rule_group_namespace" "alerting_rules" {
count = var.enable_alerting_rules ? 1 : 0
name = "accelerator-java-alerting"
workspace_id = var.managed_prometheus_workspace_id
data = <<EOF
groups:
- name: default-alert
rules:
- alert: metric:alerting_rule
expr: jvm_memory_bytes_used{job="java", area="heap"} / jvm_memory_bytes_max * 100 > 80
for: 1m
labels:
severity: warning
annotations:
summary: "JVM heap warning"
description: "JVM heap of instance `{{$labels.instance}}` from application `{{$labels.application}}` is above 80% for one minute. (current=`{{$value}}%`)"
EOF
}
resource "grafana_dashboard" "this" {
folder = var.dashboards_folder_id
config_json = file("${path.module}/dashboards/default.json")
}
@@ -1,6 +0,0 @@
apiVersion: v2
name: opentelemetry
description: A Helm chart to install otel operator
type: application
version: 0.2.0
appVersion: v0.1.0
@@ -1,29 +0,0 @@
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: otel-prometheus-role
rules:
- apiGroups:
- ""
resources:
- nodes
- nodes/proxy
- services
- endpoints
- pods
verbs:
- get
- list
- watch
- apiGroups:
- extensions
resources:
- ingresses
verbs:
- get
- list
- watch
- nonResourceURLs:
- /metrics
verbs:
- get
@@ -1,12 +0,0 @@
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
name: otel-prometheus-role-binding
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: ClusterRole
name: otel-prometheus-role
subjects:
- kind: ServiceAccount
name: adot-collector-java
namespace: adot-collector-java
@@ -1,71 +0,0 @@
apiVersion: opentelemetry.io/v1alpha1
kind: OpenTelemetryCollector
metadata:
name: adot
spec:
image: public.ecr.aws/aws-observability/aws-otel-collector:v0.22.0
mode: deployment
serviceAccount: adot-collector-java
config: |
receivers:
prometheus:
config:
global:
scrape_interval: {{ .Values.globalScrapeInterval }}
scrape_timeout: {{ .Values.globalScrapeTimeout }}
external_labels:
cluster: {{ .Values.ekscluster }}
account_id: {{ .Values.accountId }}
region: {{ .Values.region }}
scrape_configs:
- job_name: 'kubernetes-pod-jmx'
sample_limit: {{ .Values.scrapeSampleLimit }}
kubernetes_sd_configs:
- role: pod
relabel_configs:
- source_labels: [ __address__ ]
action: keep
regex: '.*:9404$'
- action: labelmap
regex: __meta_kubernetes_pod_label_(.+)
- action: replace
source_labels: [ __meta_kubernetes_namespace ]
target_label: Namespace
- source_labels: [ __meta_kubernetes_pod_name ]
action: replace
target_label: pod_name
- action: replace
source_labels: [ __meta_kubernetes_pod_container_name ]
target_label: container_name
- action: replace
source_labels: [ __meta_kubernetes_pod_controller_kind ]
target_label: pod_controller_kind
- action: replace
source_labels: [ __meta_kubernetes_pod_phase ]
target_label: pod_controller_phase
metric_relabel_configs:
- source_labels: [ __name__ ]
regex: 'jvm_gc_collection_seconds.*'
action: drop
exporters:
prometheusremotewrite:
endpoint: {{ .Values.ampurl }}
auth:
authenticator: sigv4auth
logging:
loglevel: info
extensions:
sigv4auth:
region: {{ .Values.region }}
service: "aps"
health_check:
pprof:
endpoint: :1888
zpages:
endpoint: :55679
service:
extensions: [pprof, zpages, health_check, sigv4auth]
pipelines:
metrics:
receivers: [prometheus]
exporters: [logging, prometheusremotewrite]
@@ -1,6 +0,0 @@
ampurl: ${amp_url}
region: ${region}
prometheusMetricsEndpoint: ${prometheus_metrics_endpoint}
globalScrapeInterval: ${scrape_interval}
globalScrapeTimeout: ${scrape_timeout}
scrapeSampleLimit: ${scrape_sample_limit}
-79
View File
@@ -1,79 +0,0 @@
variable "eks_cluster_id" {
description = "EKS Cluster Id"
type = string
}
variable "irsa_iam_role_path" {
description = "IAM role path for IRSA roles"
type = string
default = "/"
}
variable "irsa_iam_permissions_boundary" {
description = "IAM permissions boundary for IRSA roles"
type = string
default = null
}
variable "enable_recording_rules" {
description = "Enables or disables Managed Prometheus recording rules. Disabling this might affect some data in the dashboards"
type = bool
default = true
}
variable "enable_alerting_rules" {
description = "Enables or disables Managed Prometheus alerting rules"
type = bool
default = true
}
variable "managed_prometheus_workspace_endpoint" {
description = "Amazon Managed Prometheus Workspace Endpoint"
type = string
default = ""
}
variable "managed_prometheus_workspace_id" {
description = "Amazon Managed Prometheus Workspace ID"
type = string
default = null
}
variable "managed_prometheus_workspace_region" {
description = "Amazon Managed Prometheus Workspace's Region"
type = string
default = null
}
variable "helm_config" {
description = "Helm Config for Prometheus"
type = any
default = {}
}
variable "dashboards_folder_id" {
description = "Grafana folder ID for automatic dashboards"
type = string
}
variable "prometheus_config" {
description = "Controls default values such as scrape interval, timeouts and ports globally"
type = object({
global_scrape_interval = string
global_scrape_timeout = string
scrape_sample_limit = number
})
default = {
global_scrape_interval = "60s"
global_scrape_timeout = "15s"
scrape_sample_limit = 1000
}
nullable = false
}
variable "tags" {
description = "Additional tags (e.g. `map('BusinessUnit`,`XYZ`)"
type = map(string)
default = {}
}
-6
View File
@@ -1,6 +0,0 @@
resource "grafana_dashboard" "workloads" {
count = var.enable_dashboards ? 1 : 0
folder = var.dashboards_folder_id
config_json = file("${path.module}/dashboards/nginx.json")
}
-31
View File
@@ -1,31 +0,0 @@
data "aws_partition" "current" {}
data "aws_caller_identity" "current" {}
data "aws_region" "current" {}
data "aws_eks_cluster" "eks_cluster" {
name = var.eks_cluster_id
}
locals {
name = "adot-collector-nginx"
namespace = try(var.config.helm_config.namespace, local.name)
eks_oidc_issuer_url = replace(data.aws_eks_cluster.eks_cluster.identity[0].oidc[0].issuer, "https://", "")
eks_cluster_endpoint = data.aws_eks_cluster.eks_cluster.endpoint
context = {
aws_caller_identity_account_id = data.aws_caller_identity.current.account_id
aws_caller_identity_arn = data.aws_caller_identity.current.arn
aws_eks_cluster_endpoint = local.eks_cluster_endpoint
aws_partition_id = data.aws_partition.current.partition
aws_region_name = data.aws_region.current.name
eks_cluster_id = var.eks_cluster_id
eks_oidc_issuer_url = local.eks_oidc_issuer_url
eks_oidc_provider_arn = "arn:${data.aws_partition.current.partition}:iam::${data.aws_caller_identity.current.account_id}:oidc-provider/${local.eks_oidc_issuer_url}"
tags = var.tags
irsa_iam_role_path = var.irsa_iam_role_path
irsa_iam_permissions_boundary = var.irsa_iam_permissions_boundary
}
}
-55
View File
@@ -1,55 +0,0 @@
module "helm_addon" {
source = "github.com/aws-ia/terraform-aws-eks-blueprints//modules/kubernetes-addons/helm-addon?ref=v4.13.1"
helm_config = merge(
{
name = local.name
chart = "${path.module}/otel-config"
version = "0.2.0"
namespace = local.namespace
description = "ADOT helm Chart deployment configuration"
},
var.helm_config
)
set_values = [
{
name = "ampurl"
value = "${var.managed_prometheus_workspace_endpoint}api/v1/remote_write"
},
{
name = "region"
value = var.managed_prometheus_workspace_region
},
{
name = "prometheusMetricsEndpoint"
value = "metrics"
},
{
name = "prometheusMetricsPort"
value = 8888
},
{
name = "scrapeInterval"
value = "15s"
},
{
name = "scrapeTimeout"
value = "10s"
},
{
name = "scrapeSampleLimit"
value = 1000
}
]
irsa_config = {
create_kubernetes_namespace = try(var.helm_config["create_namespace"], true)
kubernetes_namespace = local.namespace
create_kubernetes_service_account = true
kubernetes_service_account = try(var.helm_config.service_account, local.name)
irsa_iam_policies = ["arn:${data.aws_partition.current.partition}:iam::aws:policy/AmazonPrometheusRemoteWriteAccess"]
}
addon_context = local.context
}
@@ -1,6 +0,0 @@
apiVersion: v2
name: opentelemetry
description: A Helm chart to install otel operator
type: application
version: 0.2.0
appVersion: v0.1.0
@@ -1,29 +0,0 @@
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: otel-prometheus-role
rules:
- apiGroups:
- ""
resources:
- nodes
- nodes/proxy
- services
- endpoints
- pods
verbs:
- get
- list
- watch
- apiGroups:
- extensions
resources:
- ingresses
verbs:
- get
- list
- watch
- nonResourceURLs:
- /metrics
verbs:
- get
@@ -1,12 +0,0 @@
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
name: otel-prometheus-role-binding
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: ClusterRole
name: otel-prometheus-role
subjects:
- kind: ServiceAccount
name: adot-collector-nginx
namespace: adot-collector-nginx
@@ -1,68 +0,0 @@
apiVersion: opentelemetry.io/v1alpha1
kind: OpenTelemetryCollector
metadata:
name: adot
spec:
image: public.ecr.aws/aws-observability/aws-otel-collector:v0.21.1
mode: deployment
serviceAccount: adot-collector-nginx
config: |
receivers:
prometheus:
config:
global:
scrape_interval: {{ .Values.scrapeInterval }}
scrape_timeout: {{ .Values.scrapeTimeout }}
scrape_configs:
- job_name: 'kubernetes-pod-nginx'
sample_limit: {{ .Values.scrapeSampleLimit }}
metrics_path: /{{ .Values.prometheusMetricsEndpoint }}
kubernetes_sd_configs:
- role: pod
relabel_configs:
- source_labels: [ __address__ ]
action: keep
regex: '.*:10254$'
- source_labels: [__meta_kubernetes_pod_container_name]
target_label: container
action: replace
- source_labels: [__meta_kubernetes_pod_node_name]
target_label: host
action: replace
- source_labels: [__meta_kubernetes_namespace]
target_label: namespace
action: replace
metric_relabel_configs:
- source_labels: [__name__]
regex: 'go_memstats.*'
action: drop
- source_labels: [__name__]
regex: 'go_gc.*'
action: drop
- source_labels: [__name__]
regex: 'go_threads'
action: drop
- regex: exported_host
action: labeldrop
exporters:
prometheusremotewrite:
endpoint: {{ .Values.ampurl }}
auth:
authenticator: sigv4auth
logging:
loglevel: info
extensions:
sigv4auth:
region: {{ .Values.region }}
service: "aps"
health_check:
pprof:
endpoint: :1888
zpages:
endpoint: :55679
service:
extensions: [pprof, zpages, health_check, sigv4auth]
pipelines:
metrics:
receivers: [prometheus]
exporters: [logging, prometheusremotewrite]
@@ -1,7 +0,0 @@
ampurl: ${amp_url}
region: ${region}
prometheusMetricsEndpoint: ${prometheus_metrics_endpoint}
prometheusMetricsPort: ${prometheus_metrics_port}
scrapeInterval: ${scrape_interval}
scrapeTimeout: ${scrape_timeout}
scrapeSampleLimit: ${scrape_sample_limit}
-69
View File
@@ -1,69 +0,0 @@
variable "eks_cluster_id" {
description = "EKS Cluster Id"
type = string
}
variable "helm_config" {
description = "Helm Config for Prometheus"
type = any
default = {}
}
variable "irsa_iam_role_path" {
description = "IAM role path for IRSA roles"
type = string
default = "/"
}
variable "irsa_iam_permissions_boundary" {
description = "IAM permissions boundary for IRSA roles"
type = string
default = null
}
variable "managed_prometheus_workspace_endpoint" {
description = "Amazon Managed Prometheus Workspace Endpoint"
type = string
default = ""
}
variable "managed_prometheus_workspace_id" {
description = "Amazon Managed Prometheus Workspace ID"
type = string
default = null
}
variable "managed_prometheus_workspace_region" {
description = "Amazon Managed Prometheus Workspace's Region"
type = string
default = null
}
variable "dashboards_folder_id" {
type = string
description = "Grafana folder ID for automatic dashboards"
}
variable "enable_alerting_rules" {
type = bool
default = true
description = "Enables or disables Managed Prometheus alerting rules"
}
variable "enable_dashboards" {
type = bool
description = "Enables or disables curated dashboards"
default = true
}
variable "config" {
description = "Helm Config for Prometheus"
type = any
default = {}
}
variable "tags" {
description = "Additional tags (e.g. `map('BusinessUnit`,`XYZ`)"
type = map(string)
default = {}
}