Move all dashboards to GitOps (#175)

* Typo

* Remove Grafana provider

* Temp: move dashbaords to gitOps

* Move external labels to resource attributes

* Avoid DDoS with using 0.0.0.0

* Pre-commit

* Transition in two steps

Will need to remove provider in a separate version to provide a transition path as removing this will break terraform and leave orphans in the state

* Move patterns' dashboards creation to gitOps

Standardize config objects for patterns as well

* Pre-commit

* Create AMP dashboard from external source with Grafana provider

* Fix deprecated option

* Fix Flux requirements

* Run pre-commit

* Update example with operator

* Cleanup examples

* Update multicluster example

* Update multicluster example

* Drop dead variable

* Update docs

* Change GitOps branch name

* Update docs

* Replacing Secrets Manager to SSM to store Grafana API Key (#178)

* Fixing SSM

* Fixing SSM

* Replacing Secrets Manager with SSM

* Replacing Secrets Manager with SSM

* Update architecture diagram

* Update architecture diagram

* Update README.md

* Update index.md

* Fixing Grafana Operator Version

* Fix multicluster example

* Update docs

---------

Co-authored-by: Ela AWS <51791117+elamaran11@users.noreply.github.com>
Co-authored-by: Elamaran Shanmugam <elamaran.shan@gmail.com>
This commit is contained in:
Rodrigue Koffi
2023-06-12 18:00:42 +02:00
committed by GitHub
parent c5e4c0c718
commit fa38a90efc
46 changed files with 363 additions and 4462 deletions
+7 -8
View File
@@ -67,7 +67,6 @@ See examples using this Terraform modules in the **Amazon EKS** section of [this
| Name | Description | Type | Default | Required |
|------|-------------|------|---------|:--------:|
| <a name="input_custom_metrics_config"></a> [custom\_metrics\_config](#input\_custom\_metrics\_config) | Configuration object to enable custom metrics collection | <pre>object({<br> ports = list(number)<br> # paths = optional(list(string), ["/metrics"])<br> # list of samples to be dropped by label prefix, ex: go_ -> discards go_.*<br> dropped_series_prefixes = list(string)<br> })</pre> | <pre>{<br> "dropped_series_prefixes": [<br> "unspecified"<br> ],<br> "ports": []<br>}</pre> | no |
| <a name="input_dashboards_folder_id"></a> [dashboards\_folder\_id](#input\_dashboards\_folder\_id) | Grafana folder ID for automatic dashboards | `string` | n/a | yes |
| <a name="input_eks_cluster_id"></a> [eks\_cluster\_id](#input\_eks\_cluster\_id) | EKS Cluster Id | `string` | n/a | yes |
| <a name="input_enable_alerting_rules"></a> [enable\_alerting\_rules](#input\_enable\_alerting\_rules) | Enables or disables Managed Prometheus alerting rules | `bool` | `true` | no |
| <a name="input_enable_amazon_eks_adot"></a> [enable\_amazon\_eks\_adot](#input\_enable\_amazon\_eks\_adot) | Enables the ADOT Operator on the EKS Cluster | `bool` | `true` | no |
@@ -86,11 +85,12 @@ See examples using this Terraform modules in the **Amazon EKS** section of [this
| <a name="input_enable_tracing"></a> [enable\_tracing](#input\_enable\_tracing) | (Experimental) Enables tracing with AWS X-Ray. This changes the deploy mode of the collector to daemon set. Requirement: adot add-on <= 0.58-build.0 | `bool` | `false` | no |
| <a name="input_flux_config"></a> [flux\_config](#input\_flux\_config) | FluxCD configuration | <pre>object({<br> create_namespace = bool<br> k8s_namespace = string<br> helm_chart_name = string<br> helm_chart_version = string<br> helm_release_name = string<br> helm_repo_url = string<br> helm_settings = map(string)<br> helm_values = map(any)<br> })</pre> | <pre>{<br> "create_namespace": true,<br> "helm_chart_name": "flux2",<br> "helm_chart_version": "2.7.0",<br> "helm_release_name": "observability-fluxcd-addon",<br> "helm_repo_url": "https://fluxcd-community.github.io/helm-charts",<br> "helm_settings": {},<br> "helm_values": {},<br> "k8s_namespace": "flux-system"<br>}</pre> | no |
| <a name="input_flux_gitrepository_branch"></a> [flux\_gitrepository\_branch](#input\_flux\_gitrepository\_branch) | Flux GitRepository Branch | `string` | `"main"` | no |
| <a name="input_flux_gitrepository_name"></a> [flux\_gitrepository\_name](#input\_flux\_gitrepository\_name) | Flux GitRepository name | `string` | `"aws-observability-accelerator"` | no |
| <a name="input_flux_gitrepository_url"></a> [flux\_gitrepository\_url](#input\_flux\_gitrepository\_url) | Flux GitRepository URL | `string` | `"https://github.com/aws-observability/aws-observability-accelerator"` | no |
| <a name="input_flux_kustomization_path"></a> [flux\_kustomization\_path](#input\_flux\_kustomization\_path) | Flux Kustomization Path | `string` | `"./artifacts/grafana-operator-manifests"` | no |
| <a name="input_flux_name"></a> [flux\_name](#input\_flux\_name) | Flux GitRepository and Kustomization Name | `string` | `"grafana-dashboards"` | no |
| <a name="input_go_config"></a> [go\_config](#input\_go\_config) | Grafana Operator configuration | <pre>object({<br> create_namespace = bool<br> helm_chart = string<br> helm_name = string<br> k8s_namespace = string<br> helm_release_name = string<br> helm_chart_version = string<br> })</pre> | <pre>{<br> "create_namespace": true,<br> "helm_chart": "oci://ghcr.io/grafana-operator/helm-charts/grafana-operator",<br> "helm_chart_version": "v5.0.0-rc1",<br> "helm_name": "grafana-operator",<br> "helm_release_name": "grafana-operator",<br> "k8s_namespace": "grafana-operator"<br>}</pre> | no |
| <a name="input_grafana_api_key"></a> [grafana\_api\_key](#input\_grafana\_api\_key) | Grafana API key for the Amazon Managed Grafana workspace | `string` | n/a | yes |
| <a name="input_flux_kustomization_name"></a> [flux\_kustomization\_name](#input\_flux\_kustomization\_name) | Flux Kustomization name | `string` | `"grafana-dashboards-infrastructure"` | no |
| <a name="input_flux_kustomization_path"></a> [flux\_kustomization\_path](#input\_flux\_kustomization\_path) | Flux Kustomization Path | `string` | `"./artifacts/grafana-operator-manifests/eks/infrastructure"` | no |
| <a name="input_go_config"></a> [go\_config](#input\_go\_config) | Grafana Operator configuration | <pre>object({<br> create_namespace = bool<br> helm_chart = string<br> helm_name = string<br> k8s_namespace = string<br> helm_release_name = string<br> helm_chart_version = string<br> })</pre> | <pre>{<br> "create_namespace": true,<br> "helm_chart": "oci://ghcr.io/grafana-operator/helm-charts/grafana-operator",<br> "helm_chart_version": "v5.0.0-rc3",<br> "helm_name": "grafana-operator",<br> "helm_release_name": "grafana-operator",<br> "k8s_namespace": "grafana-operator"<br>}</pre> | no |
| <a name="input_grafana_api_key"></a> [grafana\_api\_key](#input\_grafana\_api\_key) | Grafana API key for the Amazon Managed Grafana workspace | `string` | `""` | no |
| <a name="input_grafana_cluster_dashboard_url"></a> [grafana\_cluster\_dashboard\_url](#input\_grafana\_cluster\_dashboard\_url) | Dashboard URL for Cluster Grafana Dashboard JSON | `string` | `"https://raw.githubusercontent.com/aws-observability/aws-observability-accelerator/main/artifacts/grafana-dashboards/eks/infrastructure/cluster.json"` | no |
| <a name="input_grafana_kubelet_dashboard_url"></a> [grafana\_kubelet\_dashboard\_url](#input\_grafana\_kubelet\_dashboard\_url) | Dashboard URL for Kubelet Grafana Dashboard JSON | `string` | `"https://raw.githubusercontent.com/aws-observability/aws-observability-accelerator/main/artifacts/grafana-dashboards/eks/infrastructure/kubelet.json"` | no |
| <a name="input_grafana_namespace_workloads_dashboard_url"></a> [grafana\_namespace\_workloads\_dashboard\_url](#input\_grafana\_namespace\_workloads\_dashboard\_url) | Dashboard URL for Namespace Workloads Grafana Dashboard JSON | `string` | `"https://raw.githubusercontent.com/aws-observability/aws-observability-accelerator/main/artifacts/grafana-dashboards/eks/infrastructure/namespace-workloads.json"` | no |
@@ -101,14 +101,14 @@ See examples using this Terraform modules in the **Amazon EKS** section of [this
| <a name="input_helm_config"></a> [helm\_config](#input\_helm\_config) | Helm Config for Prometheus | `any` | `{}` | no |
| <a name="input_irsa_iam_permissions_boundary"></a> [irsa\_iam\_permissions\_boundary](#input\_irsa\_iam\_permissions\_boundary) | IAM permissions boundary for IRSA roles | `string` | `null` | no |
| <a name="input_irsa_iam_role_path"></a> [irsa\_iam\_role\_path](#input\_irsa\_iam\_role\_path) | IAM role path for IRSA roles | `string` | `"/"` | no |
| <a name="input_java_config"></a> [java\_config](#input\_java\_config) | Configuration object for Java/JMX monitoring | <pre>object({<br> enable_alerting_rules = bool<br> enable_recording_rules = bool<br> scrape_sample_limit = number<br> })</pre> | <pre>{<br> "enable_alerting_rules": true,<br> "enable_recording_rules": true,<br> "scrape_sample_limit": 1000<br>}</pre> | no |
| <a name="input_java_config"></a> [java\_config](#input\_java\_config) | Configuration object for Java/JMX monitoring | <pre>object({<br> enable_alerting_rules = bool<br> enable_recording_rules = bool<br> enable_dashboards = bool<br> scrape_sample_limit = number<br><br><br> flux_gitrepository_name = string<br> flux_gitrepository_url = string<br> flux_gitrepository_branch = string<br> flux_kustomization_name = string<br> flux_kustomization_path = string<br><br> grafana_dashboard_url = string<br><br> prometheus_metrics_endpoint = string<br> })</pre> | `null` | no |
| <a name="input_ksm_config"></a> [ksm\_config](#input\_ksm\_config) | Kube State metrics configuration | <pre>object({<br> create_namespace = bool<br> k8s_namespace = string<br> helm_chart_name = string<br> helm_chart_version = string<br> helm_release_name = string<br> helm_repo_url = string<br> helm_settings = map(string)<br> helm_values = map(any)<br><br> scrape_interval = string<br> scrape_timeout = string<br> })</pre> | <pre>{<br> "create_namespace": true,<br> "helm_chart_name": "kube-state-metrics",<br> "helm_chart_version": "4.24.0",<br> "helm_release_name": "kube-state-metrics",<br> "helm_repo_url": "https://prometheus-community.github.io/helm-charts",<br> "helm_settings": {},<br> "helm_values": {},<br> "k8s_namespace": "kube-system",<br> "scrape_interval": "60s",<br> "scrape_timeout": "15s"<br>}</pre> | no |
| <a name="input_logs_config"></a> [logs\_config](#input\_logs\_config) | Configuration object for logs collection | <pre>object({<br> cw_log_retention_days = number<br> })</pre> | <pre>{<br> "cw_log_retention_days": 90<br>}</pre> | no |
| <a name="input_managed_prometheus_workspace_endpoint"></a> [managed\_prometheus\_workspace\_endpoint](#input\_managed\_prometheus\_workspace\_endpoint) | Amazon Managed Prometheus Workspace Endpoint | `string` | `""` | no |
| <a name="input_managed_prometheus_workspace_id"></a> [managed\_prometheus\_workspace\_id](#input\_managed\_prometheus\_workspace\_id) | Amazon Managed Prometheus Workspace ID | `string` | `null` | no |
| <a name="input_managed_prometheus_workspace_region"></a> [managed\_prometheus\_workspace\_region](#input\_managed\_prometheus\_workspace\_region) | Amazon Managed Prometheus Workspace's Region | `string` | `null` | no |
| <a name="input_ne_config"></a> [ne\_config](#input\_ne\_config) | Node exporter configuration | <pre>object({<br> create_namespace = bool<br> k8s_namespace = string<br> helm_chart_name = string<br> helm_chart_version = string<br> helm_release_name = string<br> helm_repo_url = string<br> helm_settings = map(string)<br> helm_values = map(any)<br><br> scrape_interval = string<br> scrape_timeout = string<br> })</pre> | <pre>{<br> "create_namespace": true,<br> "helm_chart_name": "prometheus-node-exporter",<br> "helm_chart_version": "4.14.0",<br> "helm_release_name": "prometheus-node-exporter",<br> "helm_repo_url": "https://prometheus-community.github.io/helm-charts",<br> "helm_settings": {},<br> "helm_values": {},<br> "k8s_namespace": "prometheus-node-exporter",<br> "scrape_interval": "60s",<br> "scrape_timeout": "60s"<br>}</pre> | no |
| <a name="input_nginx_config"></a> [nginx\_config](#input\_nginx\_config) | Configuration object for NGINX monitoring | <pre>object({<br> enable_alerting_rules = bool<br> scrape_sample_limit = number<br> prometheus_metrics_endpoint = string<br> })</pre> | <pre>{<br> "enable_alerting_rules": true,<br> "prometheus_metrics_endpoint": "metrics",<br> "scrape_sample_limit": 1000<br>}</pre> | no |
| <a name="input_nginx_config"></a> [nginx\_config](#input\_nginx\_config) | Configuration object for NGINX monitoring | <pre>object({<br> enable_alerting_rules = bool<br> enable_recording_rules = bool<br> enable_dashboards = bool<br> scrape_sample_limit = number<br><br> flux_gitrepository_name = string<br> flux_gitrepository_url = string<br> flux_gitrepository_branch = string<br> flux_kustomization_name = string<br> flux_kustomization_path = string<br><br> grafana_dashboard_url = string<br><br> prometheus_metrics_endpoint = string<br> })</pre> | `null` | no |
| <a name="input_prometheus_config"></a> [prometheus\_config](#input\_prometheus\_config) | Controls default values such as scrape interval, timeouts and ports globally | <pre>object({<br> global_scrape_interval = string<br> global_scrape_timeout = string<br> })</pre> | <pre>{<br> "global_scrape_interval": "60s",<br> "global_scrape_timeout": "15s"<br>}</pre> | no |
| <a name="input_tags"></a> [tags](#input\_tags) | Additional tags (e.g. `map('BusinessUnit`,`XYZ`) | `map(string)` | `{}` | no |
| <a name="input_target_secret_name"></a> [target\_secret\_name](#input\_target\_secret\_name) | Target secret in Kubernetes to store the Grafana API Key Secret | `string` | `"grafana-admin-credentials"` | no |
@@ -121,5 +121,4 @@ See examples using this Terraform modules in the **Amazon EKS** section of [this
|------|-------------|
| <a name="output_eks_cluster_id"></a> [eks\_cluster\_id](#output\_eks\_cluster\_id) | EKS Cluster Id |
| <a name="output_eks_cluster_version"></a> [eks\_cluster\_version](#output\_eks\_cluster\_version) | EKS Cluster version |
| <a name="output_grafana_dashboard_urls"></a> [grafana\_dashboard\_urls](#output\_grafana\_dashboard\_urls) | URLs for dashboards created |
<!-- END OF PRE-COMMIT-TERRAFORM DOCS HOOK -->
@@ -32,8 +32,7 @@ This deploys an EKS Cluster with the External Secrets Operator. The cluster is p
|------|------|
| [aws_iam_policy.cluster_secretstore](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/resources/iam_policy) | resource |
| [aws_kms_key.secrets](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/resources/kms_key) | resource |
| [aws_secretsmanager_secret.secret](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/resources/secretsmanager_secret) | resource |
| [aws_secretsmanager_secret_version.secret](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/resources/secretsmanager_secret_version) | resource |
| [aws_ssm_parameter.secret](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/resources/ssm_parameter) | resource |
| [kubectl_manifest.cluster_secretstore](https://registry.terraform.io/providers/gavinbunney/kubectl/latest/docs/resources/manifest) | resource |
| [kubectl_manifest.secret](https://registry.terraform.io/providers/gavinbunney/kubectl/latest/docs/resources/manifest) | resource |
| [aws_region.current](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/data-sources/region) | data source |
@@ -36,12 +36,13 @@ resource "aws_iam_policy" "cluster_secretstore" {
{
"Effect": "Allow",
"Action": [
"secretsmanager:GetResourcePolicy",
"secretsmanager:GetSecretValue",
"secretsmanager:DescribeSecret",
"secretsmanager:ListSecretVersionIds"
"ssm:DescribeParameters",
"ssm:GetParameter",
"ssm:GetParameters",
"ssm:GetParametersByPath",
"ssm:GetParameterHistory"
],
"Resource": "${aws_secretsmanager_secret.secret.arn}"
"Resource": "${aws_ssm_parameter.secret.arn}"
},
{
"Effect": "Allow",
@@ -64,7 +65,7 @@ metadata:
spec:
provider:
aws:
service: SecretsManager
service: ParameterStore
region: ${data.aws_region.current.name}
auth:
jwt:
@@ -75,16 +76,15 @@ YAML
depends_on = [module.external_secrets]
}
resource "aws_secretsmanager_secret" "secret" {
recovery_window_in_days = 0
kms_key_id = aws_kms_key.secrets.arn
}
resource "aws_secretsmanager_secret_version" "secret" {
secret_id = aws_secretsmanager_secret.secret.id
secret_string = jsonencode({
resource "aws_ssm_parameter" "secret" {
name = "/terraform-accelerator/grafana-api-key"
description = "SSM Secret to store grafana API Key"
type = "SecureString"
value = jsonencode({
GF_SECURITY_ADMIN_APIKEY = var.grafana_api_key
})
key_id = aws_kms_key.secrets.id
overwrite = true
}
resource "kubectl_manifest" "secret" {
@@ -103,7 +103,7 @@ spec:
name: ${var.target_secret_name}
dataFrom:
- extract:
key: ${aws_secretsmanager_secret.secret.name}
key: ${aws_ssm_parameter.secret.name}
YAML
depends_on = [module.external_secrets]
}
+7 -6
View File
@@ -1,9 +1,11 @@
resource "kubectl_manifest" "flux_gitrepository" {
yaml_body = <<YAML
count = var.enable_dashboards ? 1 : 0
yaml_body = <<YAML
apiVersion: source.toolkit.fluxcd.io/v1beta2
kind: GitRepository
metadata:
name: ${var.flux_name}
name: ${var.flux_gitrepository_name}
namespace: flux-system
spec:
interval: 5m0s
@@ -11,9 +13,8 @@ spec:
ref:
branch: ${var.flux_gitrepository_branch}
YAML
count = var.enable_dashboards ? 1 : 0
depends_on = [module.external_secrets]
depends_on = [module.external_secrets]
}
resource "kubectl_manifest" "flux_kustomization" {
@@ -21,7 +22,7 @@ resource "kubectl_manifest" "flux_kustomization" {
apiVersion: kustomize.toolkit.fluxcd.io/v1beta2
kind: Kustomization
metadata:
name: ${var.flux_name}
name: ${var.flux_kustomization_name}
namespace: flux-system
spec:
interval: 1m0s
@@ -29,7 +30,7 @@ spec:
prune: true
sourceRef:
kind: GitRepository
name: ${var.flux_name}
name: ${var.flux_gitrepository_name}
postBuild:
substitute:
AMG_AWS_REGION: ${var.managed_prometheus_workspace_region}
+48
View File
@@ -29,4 +29,52 @@ locals {
irsa_iam_role_path = var.irsa_iam_role_path
irsa_iam_permissions_boundary = var.irsa_iam_permissions_boundary
}
java_pattern_config = {
# disabled if options from module are disabled, by default
# can be overriden by providing a config
enable_alerting_rules = var.enable_alerting_rules
enable_recording_rules = var.enable_recording_rules
enable_dashboards = var.enable_dashboards # disable flux kustomization if dashboards are disabled
scrape_sample_limit = 1000
flux_gitrepository_name = "aws-observability-accelerator"
flux_gitrepository_url = "https://github.com/aws-observability/aws-observability-accelerator"
flux_gitrepository_branch = "main"
flux_kustomization_name = "grafana-dashboards-java"
flux_kustomization_path = "./artifacts/grafana-operator-manifests/eks/java"
managed_prometheus_workspace_id = var.managed_prometheus_workspace_id
managed_prometheus_workspace_region = var.managed_prometheus_workspace_region
managed_prometheus_workspace_endpoint = var.managed_prometheus_workspace_endpoint
prometheus_metrics_endpoint = "/metrics"
grafana_url = var.grafana_url
grafana_dashboard_url = "https://raw.githubusercontent.com/aws-observability/aws-observability-accelerator/main/artifacts/grafana-dashboards/eks/java/default.json"
}
nginx_pattern_config = {
# disabled if options from module are disabled, by default
# can be overriden by providing a config
enable_alerting_rules = var.enable_alerting_rules
enable_recording_rules = var.enable_recording_rules
enable_dashboards = var.enable_dashboards
scrape_sample_limit = 1000
flux_gitrepository_name = "aws-observability-accelerator"
flux_gitrepository_url = "https://github.com/aws-observability/aws-observability-accelerator"
flux_gitrepository_branch = "main"
flux_kustomization_name = "grafana-dashboards-nginx"
flux_kustomization_path = "./artifacts/grafana-operator-manifests/eks/nginx"
managed_prometheus_workspace_id = var.managed_prometheus_workspace_id
managed_prometheus_workspace_region = var.managed_prometheus_workspace_region
managed_prometheus_workspace_endpoint = var.managed_prometheus_workspace_endpoint
prometheus_metrics_endpoint = "/metrics"
grafana_url = var.grafana_url
grafana_dashboard_url = "https://raw.githubusercontent.com/aws-observability/aws-observability-accelerator/main/artifacts/grafana-dashboards/eks/nginx/nginx.json"
}
}
+13 -14
View File
@@ -148,7 +148,11 @@ module "helm_addon" {
},
{
name = "javaScrapeSampleLimit"
value = var.java_config.scrape_sample_limit
value = try(var.java_config.scrape_sample_limit, local.java_pattern_config.scrape_sample_limit)
},
{
name = "javaPrometheusMetricsEndpoint"
value = try(var.java_config.prometheus_metrics_endpoint, local.java_pattern_config.prometheus_metrics_endpoint)
},
{
name = "enable_nginx"
@@ -156,12 +160,12 @@ module "helm_addon" {
},
{
name = "nginxScrapeSampleLimit"
value = var.nginx_config.scrape_sample_limit
value = try(var.nginx_config.scrape_sample_limit, local.nginx_pattern_config.scrape_sample_limit)
},
{
name = "nginxPrometheusMetricsEndpoint"
value = var.nginx_config.prometheus_metrics_endpoint
}
value = try(var.nginx_config.prometheus_metrics_endpoint, local.nginx_pattern_config.prometheus_metrics_endpoint)
},
]
irsa_config = {
@@ -181,23 +185,18 @@ module "helm_addon" {
}
module "java_monitoring" {
source = "./patterns/java"
count = var.enable_java ? 1 : 0
enable_dashboards = var.enable_dashboards
source = "./patterns/java"
count = var.enable_java ? 1 : 0
pattern_config = coalesce(var.java_config, local.java_pattern_config)
managed_prometheus_workspace_id = var.managed_prometheus_workspace_id
enable_alerting_rules = var.java_config.enable_alerting_rules
enable_recording_rules = var.java_config.enable_recording_rules
dashboards_folder_id = var.dashboards_folder_id
}
module "nginx_monitoring" {
source = "./patterns/nginx"
count = var.enable_nginx ? 1 : 0
managed_prometheus_workspace_id = var.managed_prometheus_workspace_id
enable_alerting_rules = var.nginx_config.enable_alerting_rules
dashboards_folder_id = var.dashboards_folder_id
pattern_config = coalesce(var.nginx_config, local.nginx_pattern_config)
}
module "fluentbit_logs" {
@@ -1701,6 +1701,7 @@ spec:
{{ if .Values.enableJava }}
- job_name: 'kubernetes-java-jmx'
sample_limit: {{ .Values.javaScrapeSampleLimit }}
metrics_path: {{ .Values.javaPrometheusMetricsEndpoint }}
kubernetes_sd_configs:
- role: pod
relabel_configs:
@@ -1733,7 +1734,7 @@ spec:
{{ if .Values.enableNginx }}
- job_name: 'kubernetes-nginx'
sample_limit: {{ .Values.nginxScrapeSampleLimit }}
metrics_path: /{{ .Values.nginxPrometheusMetricsEndpoint }}
metrics_path: {{ .Values.nginxPrometheusMetricsEndpoint }}
kubernetes_sd_configs:
- role: pod
relabel_configs:
-12
View File
@@ -1,15 +1,3 @@
output "grafana_dashboard_urls" {
value = [concat(
# grafana_dashboard.workloads[*].url,
# grafana_dashboard.nodes[*].url,
# grafana_dashboard.nsworkload[*].url,
# grafana_dashboard.kubelet[*].url,
# grafana_dashboard.cluster[*].url,
flatten(module.java_monitoring[*].grafana_dashboard_urls),
flatten(module.nginx_monitoring[*].grafana_dashboard_urls),
)]
description = "URLs for dashboards created"
}
output "eks_cluster_version" {
description = "EKS Cluster version"
value = data.aws_eks_cluster.eks_cluster.version
+4 -10
View File
@@ -22,7 +22,7 @@ Provides monitoring for Java based workloads with the following resources:
| Name | Version |
|------|---------|
| <a name="provider_aws"></a> [aws](#provider\_aws) | >= 4.0.0 |
| <a name="provider_grafana"></a> [grafana](#provider\_grafana) | >= 1.25.0 |
| <a name="provider_kubectl"></a> [kubectl](#provider\_kubectl) | >= 1.14 |
## Modules
@@ -34,21 +34,15 @@ No modules.
|------|------|
| [aws_prometheus_rule_group_namespace.alerting_rules](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/resources/prometheus_rule_group_namespace) | resource |
| [aws_prometheus_rule_group_namespace.recording_rules](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/resources/prometheus_rule_group_namespace) | resource |
| [grafana_dashboard.this](https://registry.terraform.io/providers/grafana/grafana/latest/docs/resources/dashboard) | resource |
| [kubectl_manifest.flux_kustomization](https://registry.terraform.io/providers/gavinbunney/kubectl/latest/docs/resources/manifest) | resource |
## Inputs
| Name | Description | Type | Default | Required |
|------|-------------|------|---------|:--------:|
| <a name="input_dashboards_folder_id"></a> [dashboards\_folder\_id](#input\_dashboards\_folder\_id) | Grafana folder ID for automatic dashboards | `string` | n/a | yes |
| <a name="input_enable_alerting_rules"></a> [enable\_alerting\_rules](#input\_enable\_alerting\_rules) | Enables or disables Managed Prometheus alerting rules | `bool` | `true` | no |
| <a name="input_enable_dashboards"></a> [enable\_dashboards](#input\_enable\_dashboards) | Enables or disables curated dashboards | `bool` | `true` | no |
| <a name="input_enable_recording_rules"></a> [enable\_recording\_rules](#input\_enable\_recording\_rules) | Enables or disables Managed Prometheus recording rules | `bool` | `true` | no |
| <a name="input_managed_prometheus_workspace_id"></a> [managed\_prometheus\_workspace\_id](#input\_managed\_prometheus\_workspace\_id) | Amazon Managed Prometheus Workspace ID | `string` | `null` | no |
| <a name="input_pattern_config"></a> [pattern\_config](#input\_pattern\_config) | Configuration object for Java/JMX monitoring | <pre>object({<br> enable_alerting_rules = bool<br> enable_recording_rules = bool<br> scrape_sample_limit = number<br><br> enable_recording_rules = bool<br><br> enable_dashboards = bool<br><br> flux_gitrepository_name = string<br> flux_gitrepository_url = string<br> flux_gitrepository_branch = string<br> flux_kustomization_name = string<br> flux_kustomization_path = string<br><br> managed_prometheus_workspace_id = string<br> managed_prometheus_workspace_region = string<br> managed_prometheus_workspace_endpoint = string<br><br> grafana_url = string<br> grafana_dashboard_url = string<br> })</pre> | n/a | yes |
## Outputs
| Name | Description |
|------|-------------|
| <a name="output_grafana_dashboard_urls"></a> [grafana\_dashboard\_urls](#output\_grafana\_dashboard\_urls) | URLs for dashboards created |
No outputs.
<!-- END OF PRE-COMMIT-TERRAFORM DOCS HOOK -->
File diff suppressed because it is too large Load Diff
+28 -8
View File
@@ -1,7 +1,8 @@
resource "aws_prometheus_rule_group_namespace" "recording_rules" {
count = var.enable_recording_rules ? 1 : 0
count = var.pattern_config.enable_recording_rules ? 1 : 0
name = "accelerator-java-rules"
workspace_id = var.managed_prometheus_workspace_id
workspace_id = var.pattern_config.managed_prometheus_workspace_id
data = <<EOF
groups:
- name: default-metric
@@ -12,10 +13,10 @@ EOF
}
resource "aws_prometheus_rule_group_namespace" "alerting_rules" {
count = var.enable_alerting_rules ? 1 : 0
count = var.pattern_config.enable_alerting_rules ? 1 : 0
name = "accelerator-java-alerting"
workspace_id = var.managed_prometheus_workspace_id
workspace_id = var.pattern_config.managed_prometheus_workspace_id
data = <<EOF
groups:
- name: default-alert
@@ -31,8 +32,27 @@ groups:
EOF
}
resource "grafana_dashboard" "this" {
count = var.enable_dashboards ? 1 : 0
folder = var.dashboards_folder_id
config_json = file("${path.module}/dashboards/default.json")
resource "kubectl_manifest" "flux_kustomization" {
count = var.pattern_config.enable_dashboards ? 1 : 0
yaml_body = <<YAML
apiVersion: kustomize.toolkit.fluxcd.io/v1beta2
kind: Kustomization
metadata:
name: ${var.pattern_config.flux_kustomization_name}
namespace: flux-system
spec:
interval: 1m0s
path: ${var.pattern_config.flux_kustomization_path}
prune: true
sourceRef:
kind: GitRepository
name: ${var.pattern_config.flux_gitrepository_name}
postBuild:
substitute:
AMG_AWS_REGION: ${var.pattern_config.managed_prometheus_workspace_region}
AMP_ENDPOINT_URL: ${var.pattern_config.managed_prometheus_workspace_endpoint}
AMG_ENDPOINT_URL: ${var.pattern_config.grafana_url}
GRAFANA_JAVA_JMX_DASH_URL: ${var.pattern_config.grafana_dashboard_url}
YAML
}
@@ -1,6 +0,0 @@
output "grafana_dashboard_urls" {
value = [concat(
grafana_dashboard.this[*].url,
)]
description = "URLs for dashboards created"
}
@@ -1,28 +1,26 @@
variable "enable_alerting_rules" {
description = "Enables or disables Managed Prometheus alerting rules"
type = bool
default = true
}
variable "pattern_config" {
description = "Configuration object for Java/JMX monitoring"
type = object({
enable_alerting_rules = bool
enable_recording_rules = bool
scrape_sample_limit = number
variable "enable_recording_rules" {
description = "Enables or disables Managed Prometheus recording rules"
type = bool
default = true
}
enable_recording_rules = bool
variable "managed_prometheus_workspace_id" {
description = "Amazon Managed Prometheus Workspace ID"
type = string
default = null
}
enable_dashboards = bool
variable "dashboards_folder_id" {
description = "Grafana folder ID for automatic dashboards"
type = string
}
flux_gitrepository_name = string
flux_gitrepository_url = string
flux_gitrepository_branch = string
flux_kustomization_name = string
flux_kustomization_path = string
variable "enable_dashboards" {
description = "Enables or disables curated dashboards"
type = bool
default = true
managed_prometheus_workspace_id = string
managed_prometheus_workspace_region = string
managed_prometheus_workspace_endpoint = string
grafana_url = string
grafana_dashboard_url = string
})
nullable = false
}
@@ -15,6 +15,7 @@ It provides the following resources:
| <a name="requirement_terraform"></a> [terraform](#requirement\_terraform) | >= 1.1.0 |
| <a name="requirement_aws"></a> [aws](#requirement\_aws) | >= 4.0.0 |
| <a name="requirement_grafana"></a> [grafana](#requirement\_grafana) | >= 1.25.0 |
| <a name="requirement_kubectl"></a> [kubectl](#requirement\_kubectl) | >= 1.14 |
| <a name="requirement_kubernetes"></a> [kubernetes](#requirement\_kubernetes) | >= 2.10 |
## Providers
@@ -22,7 +23,7 @@ It provides the following resources:
| Name | Version |
|------|---------|
| <a name="provider_aws"></a> [aws](#provider\_aws) | >= 4.0.0 |
| <a name="provider_grafana"></a> [grafana](#provider\_grafana) | >= 1.25.0 |
| <a name="provider_kubectl"></a> [kubectl](#provider\_kubectl) | >= 1.14 |
## Modules
@@ -33,19 +34,15 @@ No modules.
| Name | Type |
|------|------|
| [aws_prometheus_rule_group_namespace.alerting_rules](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/resources/prometheus_rule_group_namespace) | resource |
| [grafana_dashboard.workloads](https://registry.terraform.io/providers/grafana/grafana/latest/docs/resources/dashboard) | resource |
| [kubectl_manifest.flux_kustomization](https://registry.terraform.io/providers/gavinbunney/kubectl/latest/docs/resources/manifest) | resource |
## Inputs
| Name | Description | Type | Default | Required |
|------|-------------|------|---------|:--------:|
| <a name="input_dashboards_folder_id"></a> [dashboards\_folder\_id](#input\_dashboards\_folder\_id) | Grafana folder ID for automatic dashboards | `string` | n/a | yes |
| <a name="input_enable_alerting_rules"></a> [enable\_alerting\_rules](#input\_enable\_alerting\_rules) | Enables or disables Managed Prometheus alerting rules | `bool` | `true` | no |
| <a name="input_managed_prometheus_workspace_id"></a> [managed\_prometheus\_workspace\_id](#input\_managed\_prometheus\_workspace\_id) | Amazon Managed Prometheus Workspace ID | `string` | `null` | no |
| <a name="input_pattern_config"></a> [pattern\_config](#input\_pattern\_config) | Configuration object for Java/JMX monitoring | <pre>object({<br> enable_alerting_rules = bool<br> enable_recording_rules = bool<br> scrape_sample_limit = number<br><br> enable_recording_rules = bool<br><br> enable_dashboards = bool<br><br> flux_gitrepository_name = string<br> flux_gitrepository_url = string<br> flux_gitrepository_branch = string<br> flux_kustomization_name = string<br> flux_kustomization_path = string<br><br> managed_prometheus_workspace_id = string<br> managed_prometheus_workspace_region = string<br> managed_prometheus_workspace_endpoint = string<br><br> grafana_url = string<br> grafana_dashboard_url = string<br> })</pre> | n/a | yes |
## Outputs
| Name | Description |
|------|-------------|
| <a name="output_grafana_dashboard_urls"></a> [grafana\_dashboard\_urls](#output\_grafana\_dashboard\_urls) | URLs for dashboards created |
No outputs.
<!-- END OF PRE-COMMIT-TERRAFORM DOCS HOOK -->
File diff suppressed because it is too large Load Diff
+26 -5
View File
@@ -1,8 +1,8 @@
resource "aws_prometheus_rule_group_namespace" "alerting_rules" {
count = var.enable_alerting_rules ? 1 : 0
count = var.pattern_config.enable_alerting_rules ? 1 : 0
name = "accelerator-nginx-alerting"
workspace_id = var.managed_prometheus_workspace_id
workspace_id = var.pattern_config.managed_prometheus_workspace_id
data = <<EOF
groups:
- name: Nginx-HTTP-4xx-error-rate
@@ -37,8 +37,29 @@ groups:
description: "Nginx p99 latency is higher than 3 seconds\n VALUE = {{ $value }}\n LABELS = {{ $labels }}"
EOF
}
resource "grafana_dashboard" "workloads" {
folder = var.dashboards_folder_id
config_json = file("${path.module}/dashboards/nginx.json")
resource "kubectl_manifest" "flux_kustomization" {
count = var.pattern_config.enable_dashboards ? 1 : 0
yaml_body = <<YAML
apiVersion: kustomize.toolkit.fluxcd.io/v1beta2
kind: Kustomization
metadata:
name: ${var.pattern_config.flux_kustomization_name}
namespace: flux-system
spec:
interval: 1m0s
path: ${var.pattern_config.flux_kustomization_path}
prune: true
sourceRef:
kind: GitRepository
name: ${var.pattern_config.flux_gitrepository_name}
postBuild:
substitute:
AMG_AWS_REGION: ${var.pattern_config.managed_prometheus_workspace_region}
AMP_ENDPOINT_URL: ${var.pattern_config.managed_prometheus_workspace_endpoint}
AMG_ENDPOINT_URL: ${var.pattern_config.grafana_url}
GRAFANA_NGINX_DASH_URL: ${var.pattern_config.grafana_dashboard_url}
YAML
}
@@ -1,4 +0,0 @@
output "grafana_dashboard_urls" {
value = [concat(grafana_dashboard.workloads[*].url)]
description = "URLs for dashboards created"
}
@@ -1,16 +1,26 @@
variable "managed_prometheus_workspace_id" {
description = "Amazon Managed Prometheus Workspace ID"
type = string
default = null
}
variable "pattern_config" {
description = "Configuration object for Java/JMX monitoring"
type = object({
enable_alerting_rules = bool
enable_recording_rules = bool
scrape_sample_limit = number
variable "dashboards_folder_id" {
type = string
description = "Grafana folder ID for automatic dashboards"
}
enable_recording_rules = bool
variable "enable_alerting_rules" {
type = bool
default = true
description = "Enables or disables Managed Prometheus alerting rules"
enable_dashboards = bool
flux_gitrepository_name = string
flux_gitrepository_url = string
flux_gitrepository_branch = string
flux_kustomization_name = string
flux_kustomization_path = string
managed_prometheus_workspace_id = string
managed_prometheus_workspace_region = string
managed_prometheus_workspace_endpoint = string
grafana_url = string
grafana_dashboard_url = string
})
nullable = false
}
@@ -10,6 +10,10 @@ terraform {
source = "hashicorp/kubernetes"
version = ">= 2.10"
}
kubectl = {
source = "gavinbunney/kubectl"
version = ">= 1.14"
}
grafana = {
source = "grafana/grafana"
version = ">= 1.25.0"
+42 -22
View File
@@ -51,11 +51,6 @@ variable "managed_prometheus_workspace_region" {
default = null
}
variable "dashboards_folder_id" {
description = "Grafana folder ID for automatic dashboards"
type = string
}
variable "enable_alerting_rules" {
description = "Enables or disables Managed Prometheus alerting rules"
type = bool
@@ -74,10 +69,16 @@ variable "enable_dashboards" {
default = true
}
variable "flux_name" {
description = "Flux GitRepository and Kustomization Name"
variable "flux_kustomization_name" {
description = "Flux Kustomization name"
type = string
default = "grafana-dashboards"
default = "grafana-dashboards-infrastructure"
}
variable "flux_gitrepository_name" {
description = "Flux GitRepository name"
type = string
default = "aws-observability-accelerator"
}
variable "flux_gitrepository_url" {
@@ -95,7 +96,7 @@ variable "flux_gitrepository_branch" {
variable "flux_kustomization_path" {
description = "Flux Kustomization Path"
type = string
default = "./artifacts/grafana-operator-manifests"
default = "./artifacts/grafana-operator-manifests/eks/infrastructure"
}
variable "enable_kube_state_metrics" {
@@ -249,14 +250,23 @@ variable "java_config" {
type = object({
enable_alerting_rules = bool
enable_recording_rules = bool
enable_dashboards = bool
scrape_sample_limit = number
flux_gitrepository_name = string
flux_gitrepository_url = string
flux_gitrepository_branch = string
flux_kustomization_name = string
flux_kustomization_path = string
grafana_dashboard_url = string
prometheus_metrics_endpoint = string
})
default = {
enable_alerting_rules = true
enable_recording_rules = true
scrape_sample_limit = 1000
}
# defaults are pre-computed in locals.tf, provide a full definition to override
default = null
}
variable "enable_nginx" {
@@ -265,19 +275,28 @@ variable "enable_nginx" {
default = false
}
variable "nginx_config" {
description = "Configuration object for NGINX monitoring"
type = object({
enable_alerting_rules = bool
scrape_sample_limit = number
enable_alerting_rules = bool
enable_recording_rules = bool
enable_dashboards = bool
scrape_sample_limit = number
flux_gitrepository_name = string
flux_gitrepository_url = string
flux_gitrepository_branch = string
flux_kustomization_name = string
flux_kustomization_path = string
grafana_dashboard_url = string
prometheus_metrics_endpoint = string
})
default = {
enable_alerting_rules = true
scrape_sample_limit = 1000
prometheus_metrics_endpoint = "metrics"
}
# defaults are pre-computed in locals.tf, provide a full definition to override
default = null
}
variable "enable_logs" {
@@ -353,7 +372,7 @@ variable "go_config" {
helm_name = "grafana-operator"
k8s_namespace = "grafana-operator"
helm_release_name = "grafana-operator"
helm_chart_version = "v5.0.0-rc1"
helm_chart_version = "v5.0.0-rc3"
}
nullable = false
}
@@ -367,6 +386,7 @@ variable "enable_external_secrets" {
variable "grafana_api_key" {
description = "Grafana API key for the Amazon Managed Grafana workspace"
type = string
default = ""
}
variable "grafana_url" {
@@ -1,795 +0,0 @@
{
"annotations": {
"list": [
{
"builtIn": 1,
"datasource": "-- Grafana --",
"enable": true,
"hide": true,
"iconColor": "rgba(0, 211, 255, 1)",
"name": "Annotations & Alerts",
"target": {
"limit": 100,
"matchAny": false,
"tags": [],
"type": "dashboard"
},
"type": "dashboard"
}
]
},
"description": "Dashboard for Amazon Managed Prometheus",
"editable": true,
"fiscalYearStartMonth": 0,
"graphTooltip": 0,
"id": 51,
"iteration": 1666292684202,
"links": [],
"liveNow": false,
"panels": [
{
"gridPos": {
"h": 7,
"w": 5,
"x": 0,
"y": 0
},
"id": 16,
"options": {
"content": "# Ingestion Usage Metrics\n\nMetrics relating to ingestion usage of the AMP service",
"mode": "markdown"
},
"pluginVersion": "8.4.7",
"title": "Usage",
"type": "text"
},
{
"datasource": {
"type": "cloudwatch",
"uid": "$datasource"
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"custom": {
"axisLabel": "",
"axisPlacement": "auto",
"barAlignment": 0,
"drawStyle": "line",
"fillOpacity": 0,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 5,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "auto",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
},
{
"color": "red",
"value": 80
}
]
}
},
"overrides": []
},
"gridPos": {
"h": 7,
"w": 9,
"x": 5,
"y": 0
},
"id": 6,
"options": {
"legend": {
"calcs": [],
"displayMode": "list",
"placement": "bottom"
},
"tooltip": {
"mode": "single",
"sort": "none"
}
},
"targets": [
{
"alias": "",
"datasource": {
"type": "cloudwatch",
"uid": "$datasource"
},
"dimensions": {},
"expression": "SELECT SUM(ResourceCount) FROM SCHEMA(\"AWS/Usage\", Class,Resource,ResourceId,Service,Type) WHERE Type = 'Resource' AND ResourceId = '$WorkspaceID' AND Resource = 'ActiveSeries' AND Service = 'Prometheus' AND Class = 'None'",
"id": "",
"matchExact": true,
"metricEditorMode": 1,
"metricName": "",
"metricQueryType": 0,
"namespace": "",
"period": "",
"queryMode": "Metrics",
"refId": "A",
"region": "default",
"sqlExpression": "",
"statistic": "Average"
}
],
"title": "Active Series Metrics",
"type": "timeseries"
},
{
"datasource": {
"type": "cloudwatch",
"uid": "$datasource"
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"custom": {
"axisLabel": "",
"axisPlacement": "auto",
"barAlignment": 0,
"drawStyle": "line",
"fillOpacity": 0,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 5,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "auto",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
},
{
"color": "red",
"value": 80
}
]
}
},
"overrides": []
},
"gridPos": {
"h": 7,
"w": 9,
"x": 14,
"y": 0
},
"id": 2,
"options": {
"legend": {
"calcs": [],
"displayMode": "list",
"placement": "bottom"
},
"tooltip": {
"mode": "single",
"sort": "none"
}
},
"targets": [
{
"alias": "",
"datasource": {
"type": "cloudwatch",
"uid": "$datasource"
},
"dimensions": {},
"expression": "SELECT AVG(ResourceCount) FROM SCHEMA(\"AWS/Usage\", Class,Resource,ResourceId,Service,Type) WHERE Type = 'Resource' AND ResourceId = '$WorkspaceID' AND Resource = 'IngestionRate' AND Service = 'Prometheus' AND Class = 'None'",
"id": "",
"matchExact": true,
"metricEditorMode": 1,
"metricName": "",
"metricQueryType": 0,
"namespace": "",
"period": "",
"queryMode": "Metrics",
"refId": "A",
"region": "default",
"sqlExpression": "",
"statistic": "Average"
}
],
"title": "Workspace Ingestion Rate",
"type": "timeseries"
},
{
"gridPos": {
"h": 8,
"w": 5,
"x": 0,
"y": 7
},
"id": 22,
"options": {
"content": "# Billing\n\nContains information relating to the cost of AMP\n\n",
"mode": "markdown"
},
"pluginVersion": "8.4.7",
"title": "Billing",
"type": "text"
},
{
"datasource": {
"type": "cloudwatch",
"uid": "$datasource"
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"custom": {
"axisLabel": "",
"axisPlacement": "auto",
"barAlignment": 0,
"drawStyle": "line",
"fillOpacity": 0,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 5,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "auto",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
},
{
"color": "red",
"value": 80
}
]
}
},
"overrides": []
},
"gridPos": {
"h": 8,
"w": 18,
"x": 5,
"y": 7
},
"id": 24,
"options": {
"legend": {
"calcs": [],
"displayMode": "list",
"placement": "bottom"
},
"tooltip": {
"mode": "single",
"sort": "none"
}
},
"targets": [
{
"alias": "",
"datasource": {
"type": "cloudwatch",
"uid": "$datasource"
},
"dimensions": {},
"expression": "SELECT SUM(EstimatedCharges) FROM SCHEMA(\"AWS/Billing\", Currency,ServiceName) WHERE ServiceName = 'AmazonPrometheus'",
"id": "",
"matchExact": true,
"metricEditorMode": 1,
"metricName": "",
"metricQueryType": 0,
"namespace": "",
"period": "",
"queryMode": "Metrics",
"refId": "A",
"region": "default",
"sqlExpression": "",
"statistic": "Average"
}
],
"title": "Sum of Estimated AMP Charges (total)",
"type": "timeseries"
},
{
"gridPos": {
"h": 9,
"w": 5,
"x": 0,
"y": 15
},
"id": 14,
"options": {
"content": "# Alert Usage Metrics\n\nMetrics associated with Alertmanager Alert Usage",
"mode": "markdown"
},
"pluginVersion": "8.4.7",
"title": "Alerts",
"type": "text"
},
{
"datasource": {
"type": "cloudwatch",
"uid": "$datasource"
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "thresholds"
},
"mappings": [],
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
},
{
"color": "red",
"value": 1
}
]
}
},
"overrides": []
},
"gridPos": {
"h": 9,
"w": 4,
"x": 5,
"y": 15
},
"id": 4,
"options": {
"colorMode": "value",
"graphMode": "area",
"justifyMode": "auto",
"orientation": "auto",
"reduceOptions": {
"calcs": [
"lastNotNull"
],
"fields": "",
"values": false
},
"textMode": "auto"
},
"pluginVersion": "8.4.7",
"targets": [
{
"alias": "",
"datasource": {
"type": "cloudwatch",
"uid": "$datasource"
},
"dimensions": {},
"expression": "SELECT AVG(ResourceCount) FROM SCHEMA(\"AWS/Usage\", Class,Resource,ResourceId,Service,Type) WHERE Type = 'Resource' AND ResourceId = '$WorkspaceID' AND Resource = 'ActiveAlerts' AND Service = 'Prometheus' AND Class = 'None'",
"id": "",
"matchExact": true,
"metricEditorMode": 1,
"metricName": "",
"metricQueryType": 0,
"namespace": "",
"period": "",
"queryMode": "Metrics",
"refId": "A",
"region": "default",
"sqlExpression": "",
"statistic": "Average"
}
],
"title": "Active Alerts",
"type": "stat"
},
{
"datasource": {
"type": "cloudwatch",
"uid": "$datasource"
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "thresholds"
},
"mappings": [],
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
},
{
"color": "red",
"value": 1
}
]
}
},
"overrides": []
},
"gridPos": {
"h": 9,
"w": 5,
"x": 9,
"y": 15
},
"id": 12,
"options": {
"colorMode": "value",
"graphMode": "area",
"justifyMode": "auto",
"orientation": "auto",
"reduceOptions": {
"calcs": [
"lastNotNull"
],
"fields": "",
"values": false
},
"textMode": "auto"
},
"pluginVersion": "8.4.7",
"targets": [
{
"alias": "",
"datasource": {
"type": "cloudwatch",
"uid": "$datasource"
},
"dimensions": {},
"expression": "SELECT AVG(AlertManagerNotificationsFailed) FROM SCHEMA(\"AWS/Prometheus\", Workspace) WHERE Workspace = '$WorkspaceID'",
"id": "",
"matchExact": true,
"metricEditorMode": 1,
"metricName": "",
"metricQueryType": 0,
"namespace": "",
"period": "",
"queryMode": "Metrics",
"refId": "A",
"region": "default",
"sqlExpression": "",
"statistic": "Average"
}
],
"title": "Alert Manager Notifications Failed",
"type": "stat"
},
{
"datasource": {
"type": "cloudwatch",
"uid": "$datasource"
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "thresholds"
},
"mappings": [],
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"gridPos": {
"h": 9,
"w": 4,
"x": 14,
"y": 15
},
"id": 10,
"options": {
"colorMode": "value",
"graphMode": "area",
"justifyMode": "auto",
"orientation": "auto",
"reduceOptions": {
"calcs": [
"lastNotNull"
],
"fields": "",
"values": false
},
"textMode": "auto"
},
"pluginVersion": "8.4.7",
"targets": [
{
"alias": "",
"datasource": {
"type": "cloudwatch",
"uid": "$datasource"
},
"dimensions": {},
"expression": "SELECT AVG(AlertManagerAlertsReceived) FROM SCHEMA(\"AWS/Prometheus\", Workspace) WHERE Workspace = '$WorkspaceID'",
"id": "",
"matchExact": true,
"metricEditorMode": 1,
"metricName": "",
"metricQueryType": 0,
"namespace": "",
"period": "",
"queryMode": "Metrics",
"refId": "A",
"region": "default",
"sqlExpression": "",
"statistic": "Average"
}
],
"title": "Alert Manager Alerts Received",
"type": "stat"
},
{
"datasource": {
"type": "cloudwatch",
"uid": "$datasource"
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "thresholds"
},
"mappings": [],
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
},
{
"color": "red",
"value": 80
}
]
}
},
"overrides": []
},
"gridPos": {
"h": 9,
"w": 5,
"x": 18,
"y": 15
},
"id": 8,
"options": {
"colorMode": "value",
"graphMode": "area",
"justifyMode": "auto",
"orientation": "auto",
"reduceOptions": {
"calcs": [
"lastNotNull"
],
"fields": "",
"values": false
},
"textMode": "auto"
},
"pluginVersion": "8.4.7",
"targets": [
{
"alias": "",
"datasource": {
"type": "cloudwatch",
"uid": "$datasource"
},
"dimensions": {},
"expression": "SELECT AVG(ResourceCount) FROM SCHEMA(\"AWS/Usage\", Class,Resource,ResourceId,Service,Type) WHERE Type = 'Resource' AND ResourceId = '$WorkspaceID' AND Resource = 'SizeOfAlerts' AND Service = 'Prometheus' AND Class = 'None'",
"id": "",
"matchExact": true,
"metricEditorMode": 1,
"metricName": "",
"metricQueryType": 0,
"namespace": "",
"period": "",
"queryMode": "Metrics",
"refId": "A",
"region": "default",
"sqlExpression": "",
"statistic": "Average"
}
],
"title": "Size of Alerts",
"type": "stat"
},
{
"gridPos": {
"h": 7,
"w": 5,
"x": 0,
"y": 24
},
"id": 20,
"options": {
"content": "# AMP Vended Logs\n\nLast 25 log events from AMP Vended Logs for alert and rule evaluation",
"mode": "markdown"
},
"pluginVersion": "8.4.7",
"title": "AMP Logs",
"type": "text"
},
{
"datasource": {
"type": "cloudwatch",
"uid": "$datasource"
},
"gridPos": {
"h": 7,
"w": 18,
"x": 5,
"y": 24
},
"id": 18,
"options": {
"dedupStrategy": "none",
"enableLogDetails": true,
"prettifyLogMessage": false,
"showCommonLabels": false,
"showLabels": false,
"showTime": false,
"sortOrder": "Descending",
"wrapLogMessage": false
},
"targets": [
{
"datasource": {
"type": "cloudwatch",
"uid": "$datasource"
},
"expression": "fields @timestamp, @message\n| sort @timestamp desc\n| limit 25",
"id": "",
"logGroupNames": [
"/aws/vendedlogs/amp"
],
"namespace": "",
"queryMode": "Logs",
"refId": "A",
"region": "default",
"statsGroups": []
}
],
"timeFrom": "6h",
"timeShift": "6h",
"title": "AMP Vended Logs",
"type": "logs"
}
],
"refresh": "",
"schemaVersion": 35,
"style": "dark",
"tags": [],
"templating": {
"list": [
{
"current": {
"selected": true,
"text": [
"ws-e8b003eb-0528-4208-b31c-edf4598d5f66"
],
"value": [
"ws-e8b003eb-0528-4208-b31c-edf4598d5f66"
]
},
"datasource": {
"type": "cloudwatch",
"uid": "$datasource"
},
"definition": "dimension_values(default,AWS/Prometheus,RuleEvaluations,Workspace)",
"hide": 0,
"includeAll": false,
"multi": true,
"name": "WorkspaceID",
"options": [],
"query": "dimension_values(default,AWS/Prometheus,RuleEvaluations,Workspace)",
"refresh": 1,
"regex": "",
"skipUrlSync": false,
"sort": 0,
"type": "query"
},
{
"current": {
"selected": false,
"text": "Amazon CloudWatch us-west-2",
"value": "Amazon CloudWatch us-west-2"
},
"hide": 0,
"includeAll": false,
"multi": false,
"name": "datasource",
"options": [],
"query": "cloudwatch",
"refresh": 1,
"regex": "",
"skipUrlSync": false,
"type": "datasource"
}
]
},
"time": {
"from": "now-6h",
"to": "now"
},
"timepicker": {},
"timezone": "",
"title": "AMP Accelerator Dashboard",
"uid": "",
"version": 1,
"weekStart": ""
}
+10 -2
View File
@@ -14,17 +14,25 @@ resource "grafana_data_source" "cloudwatch" {
# Giving priority to Managed Prometheus datasources
is_default = false
json_data {
json_data_encoded = jsonencode({
default_region = var.aws_region
sigv4_auth = true
sigv4_auth_type = "workspace-iam-role"
sigv4_region = var.aws_region
})
}
data "http" "dashboard" {
url = "https://raw.githubusercontent.com/aws-observability/aws-observability-accelerator/a72787328e493c4628680487e3c885fc395d1c56/artifacts/grafana-dashboards/amp/amp-dashboard.json"
request_headers = {
Accept = "application/json"
}
}
resource "grafana_dashboard" "this" {
folder = var.dashboards_folder_id
config_json = file("${path.module}/dashboards/amp-dashboard.json")
config_json = data.http.dashboard.response_body
}
module "billing" {
@@ -1,8 +1,3 @@
variable "dashboards_folder_id" {
description = "Grafana folder ID for automatic dashboards"
type = string
}
variable "aws_region" {
description = "AWS Region"
type = string
@@ -24,3 +19,9 @@ variable "ingestion_rate_threshold" {
type = number
default = 70000
}
variable "dashboards_folder_id" {
description = "Grafana folder ID for automatic dashboards"
default = "0"
type = string
}
@@ -6,9 +6,17 @@ terraform {
source = "hashicorp/aws"
version = ">= 4.0.0"
}
helm = {
source = "hashicorp/helm"
version = ">= 2.4.1"
}
grafana = {
source = "grafana/grafana"
version = ">= 1.25.0"
}
http = {
source = "hashicorp/http"
version = ">= 3.3.0"
}
}
}