Files
terraform-aws-observability…/docs/eks.md
T
Kevin Lewin 0f60fb8a8d Docs (#69)
* working first run

* removing core module dependencies

* adding CW datasource

* alarms MVP

* readmes

* Adding Screenshot

* Adding billing note

* adding billing module

* Revert "adding billing module"

This reverts commit 40d667e37db1036cd71a471ef2fde83ec02aaa13.

reverting

* adding billing module

* Updating Screenshot

* resolving feedback

* removing unused modules

* fmt

* Support for tf 1.3.x

* removing unused variables

* support alarms for multiple workspaces

* Updating Readme

* docs first draft

* indigo nav fix

* Simplify docs

* RUM White Logo

* amp docs

* removing billing docs

* Change docs structure, reword infrastructure monitoring doc

* Pre-commit fixes

* drop dead code

* Add concepts page

* Update contributors

* Update java

* Update docs

* Docs for workloads

* Update docs site

* typos

* pre-commit fixes

* Update pre-commit

* Update pre-commit

Co-authored-by: Rodrigue Koffi <bonclay7@users.noreply.github.com>
2023-01-09 17:03:08 -05:00

6.1 KiB

Amazon EKS cluster monitoring

This example demonstrates how to monitor your Amazon Elastic Kubernetes Service (Amazon EKS) cluster with the Observability Accelerator's EKS infrastructure module.

Monitoring Amazon Elastic Kubernetes Service (Amazon EKS) has two categories: the control plane and the Amazon EKS nodes (with Kubernetes objects). The Amazon EKS control plane consists of control plane nodes that run the Kubernetes software, such as etcd and the Kubernetes API server. To read more on the components of an Amazon EKS cluster, please read the service documentation.

The Amazon EKS infrastructure Terraform modules focuses on metrics collection to Amazon Managed Service for Prometheus using the AWS Distro for OpenTelemetry Operator for Amazon EKS. Additionally, it provides default dashboards to get a comprehensible visibility on the nodes, namespaces, pods, and kubelet operations health. Finally, you get curated Prometheus recording rules and alerts to operate your cluster.

Prerequisites

Make sure to complete the prerequisites section before proceeding.

Setup

1. Download sources and initialize Terraform

git clone https://github.com/aws-observability/terraform-aws-observability-accelerator.git
cd examples/existing-cluster-with-base-and-infra
terraform init

2. AWS Region

Specify the AWS Region where the resources will be deployed:

export TF_VAR_aws_region=xxx

3. Amazon EKS Cluster

To run this example, you need to provide your EKS cluster name. If you don't have a cluster ready, visit this example first to create a new one.

Specify your cluster name:

export TF_VAR_eks_cluster_id=xxx

4. Amazon Managed Service for Prometheus workspace (optional)

By default, we create an Amazon Managed Service for Prometheus workspace for you. However, if you have an existing workspace you want to reuse, edit and run:

export TF_VAR_managed_prometheus_workspace_id=ws-xxx

To create a workspace outside of Terraform's state, simply run:

aws amp create-workspace --alias observability-accelerator --query '.workspaceId' --output text

5. Amazon Managed Grafana workspace

To run this example you need an Amazon Managed Grafana workspace. If you have an existing workspace, edit and run:

export TF_VAR_managed_grafana_workspace_id=g-xxx

To create a new one, within this example's Terraform state (sharing the same lifecycle with all the other resources created by Terraform):

  • Edit main.tf and set enable_managed_grafana = true
  • Run
terraform init
terraform apply -target "module.eks_observability_accelerator.module.managed_grafana[0].aws_grafana_workspace.this[0]"
export TF_VAR_managed_grafana_workspace_id=$(terraform output --raw managed_grafana_workspace_id)

6. Grafana API Key

Amazon Managed Grafana provides a control plane API for generating Grafana API keys. As a security best practice, we will provide to Terraform a short lived API key to run the apply or destroy command.

Ensure you have necessary IAM permissions (CreateWorkspaceApiKey, DeleteWorkspaceApiKey)

export TF_VAR_grafana_api_key=`aws grafana create-workspace-api-key --key-name "observability-accelerator-$(date +%s)" --key-role ADMIN --seconds-to-live 1200 --workspace-id $TF_VAR_managed_grafana_workspace_id --query key --output text`

Deploy

Simply run this command to deploy the example

terraform apply

Visualization

  1. Prometheus datasource on Grafana

Open your Grafana workspace and under Configuration -> Data sources, you should see aws-observability-accelerator. Open and click Save & test. You should see a notification confirming that the Amazon Managed Service for Prometheus workspace is ready to be used on Grafana.

  1. Grafana dashboards

Go to the Dashboards panel of your Grafana workspace. You should see a list of dashboards under the Observability Accelerator Dashboards

image

Open a specific dashboard and you should be able to view its visualization

cluster headlines
  1. Amazon Managed Service for Prometheus rules and alerts

Open the Amazon Managed Service for Prometheus console and view the details of your workspace. Under the Rules management tab, you should find new rules deployed.

image

To setup your alert receiver, with Amazon SNS, follow this documentation

Destroy resources

If you leave this stack running, you will continue to incur charges. To remove all resources created by Terraform, refresh your Grafana API key and run the command below.

Be careful, this command will removing everything created by Terraform. If you wish to keep your Amazon Managed Grafana or Amazon Managed Service for Prometheus workspaces. Remove them from your terraform state before running the destroy command.

terraform destroy

To remove resources from your Terraform state, run

# grafana workspace
terraform state rm "module.eks_observability_accelerator.module.managed_grafana[0].aws_grafana_workspace.this[0]"

# prometheus workspace
terraform state rm "module.eks_observability_accelerator.aws_prometheus_workspace.this[0]"

Note: To view all the features proposed by this module, visit the module documentation.