Compose EKS monitoring modules (#115)

* Move modules around

* Update amp billing source

* Merge Java monitoring to EKS

* Update docs

* Merge nginx pattern

* Pre-commit

* Add save and test URL output

* Move EKS dependencies to EKS monitoring module

* update docs

* Update examples and docs

* Add java doc

* Add NGINX doc

* Update nginx doc

* Fix amp monitoring example path

* Fix pre-commit

* Todo: move to main after merge

* Update docs, fix tags
This commit is contained in:
Rodrigue Koffi
2023-02-20 18:36:08 +01:00
committed by GitHub
parent fe83579997
commit daed34db80
99 changed files with 887 additions and 1477 deletions
+14 -16
View File
@@ -28,6 +28,15 @@ visit the [Amazon EKS cluster monitoring documentation](https://aws-observabilit
The sections below demonstrate how you can leverage AWS Observability Accelerator The sections below demonstrate how you can leverage AWS Observability Accelerator
to enable monitoring to an existing EKS cluster. to enable monitoring to an existing EKS cluster.
### v2.x changes
v2+ releases introduces couple of breaking changes compared to previous versions:
- `modules/workloads/infra` module moves to `modules/eks-monitoring`
- All EKS configuration options moves from the base module to the `eks-monitoring` module
- All EKS workload modules `modules/workloads/{java,nginx}` merge into `eks-monitoring` as configuration options (patterns), see [examples](./examples) to provide a more complete visiblity.
- All examples have been updated to reflect these changes
### Base Module ### Base Module
The base module allows you to configure the AWS Observability services for your cluster and The base module allows you to configure the AWS Observability services for your cluster and
@@ -38,7 +47,7 @@ and ADOT Operator deployed for you and ready to receive your data.
The base module serve as an anchor to the workload modules and cannot run on its own. The base module serve as an anchor to the workload modules and cannot run on its own.
```hcl ```hcl
module "eks_observability_accelerator" { module "aws_observability_accelerator" {
# use release tags and check for the latest versions # use release tags and check for the latest versions
# https://github.com/aws-observability/terraform-aws-observability-accelerator/releases # https://github.com/aws-observability/terraform-aws-observability-accelerator/releases
source = "github.com/aws-observability/terraform-aws-observability-accelerator?ref=v1.6.1" source = "github.com/aws-observability/terraform-aws-observability-accelerator?ref=v1.6.1"
@@ -46,7 +55,7 @@ module "eks_observability_accelerator" {
aws_region = "eu-west-1" aws_region = "eu-west-1"
eks_cluster_id = "my-eks-cluster" eks_cluster_id = "my-eks-cluster"
# As Grafana shares a different lifecycle, it's best to use an existing workspace. # As Grafana shares a different lifecycle, we recommend using an existing workspace.
managed_grafana_workspace_id = var.managed_grafana_workspace_id managed_grafana_workspace_id = var.managed_grafana_workspace_id
grafana_api_key = var.grafana_api_key grafana_api_key = var.grafana_api_key
} }
@@ -55,7 +64,7 @@ module "eks_observability_accelerator" {
You can optionally reuse an existing Amazon Managed Servce for Prometheus Workspace: You can optionally reuse an existing Amazon Managed Servce for Prometheus Workspace:
```hcl ```hcl
module "eks_observability_accelerator" { module "aws_observability_accelerator" {
# use release tags and check for the latest versions # use release tags and check for the latest versions
# https://github.com/aws-observability/terraform-aws-observability-accelerator/releases # https://github.com/aws-observability/terraform-aws-observability-accelerator/releases
source = "github.com/aws-observability/terraform-aws-observability-accelerator?ref=v1.6.1" source = "github.com/aws-observability/terraform-aws-observability-accelerator?ref=v1.6.1"
@@ -78,10 +87,9 @@ View all the configuration options in the module documentation below.
### Workload modules ### Workload modules
[Workloads modules](./modules/workloads) are provided, which essentially provide curated [Workloads modules](./modules) are provided, which essentially provide curated
metrics collection, alerting rule and Grafana dashboards. metrics collection, alerting rule and Grafana dashboards.
#### Infrastructure monitoring #### Infrastructure monitoring
```hcl ```hcl
@@ -143,7 +151,6 @@ If you are interested in contributing, see the [Contribution guide](https://gith
| Name | Source | Version | | Name | Source | Version |
|------|--------|---------| |------|--------|---------|
| <a name="module_managed_grafana"></a> [managed\_grafana](#module\_managed\_grafana) | terraform-aws-modules/managed-service-grafana/aws | ~> 1.3 | | <a name="module_managed_grafana"></a> [managed\_grafana](#module\_managed\_grafana) | terraform-aws-modules/managed-service-grafana/aws | ~> 1.3 |
| <a name="module_operator"></a> [operator](#module\_operator) | ./modules/add-ons/adot-operator | n/a |
## Resources ## Resources
@@ -153,10 +160,7 @@ If you are interested in contributing, see the [Contribution guide](https://gith
| [aws_prometheus_workspace.this](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/resources/prometheus_workspace) | resource | | [aws_prometheus_workspace.this](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/resources/prometheus_workspace) | resource |
| [grafana_data_source.amp](https://registry.terraform.io/providers/grafana/grafana/1.25.0/docs/resources/data_source) | resource | | [grafana_data_source.amp](https://registry.terraform.io/providers/grafana/grafana/1.25.0/docs/resources/data_source) | resource |
| [grafana_folder.this](https://registry.terraform.io/providers/grafana/grafana/1.25.0/docs/resources/folder) | resource | | [grafana_folder.this](https://registry.terraform.io/providers/grafana/grafana/1.25.0/docs/resources/folder) | resource |
| [aws_caller_identity.current](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/data-sources/caller_identity) | data source |
| [aws_eks_cluster.eks_cluster](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/data-sources/eks_cluster) | data source |
| [aws_grafana_workspace.this](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/data-sources/grafana_workspace) | data source | | [aws_grafana_workspace.this](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/data-sources/grafana_workspace) | data source |
| [aws_partition.current](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/data-sources/partition) | data source |
| [aws_region.current](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/data-sources/region) | data source | | [aws_region.current](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/data-sources/region) | data source |
## Inputs ## Inputs
@@ -164,15 +168,10 @@ If you are interested in contributing, see the [Contribution guide](https://gith
| Name | Description | Type | Default | Required | | Name | Description | Type | Default | Required |
|------|-------------|------|---------|:--------:| |------|-------------|------|---------|:--------:|
| <a name="input_aws_region"></a> [aws\_region](#input\_aws\_region) | AWS Region | `string` | n/a | yes | | <a name="input_aws_region"></a> [aws\_region](#input\_aws\_region) | AWS Region | `string` | n/a | yes |
| <a name="input_eks_cluster_id"></a> [eks\_cluster\_id](#input\_eks\_cluster\_id) | Name of the EKS cluster | `string` | n/a | yes |
| <a name="input_enable_alertmanager"></a> [enable\_alertmanager](#input\_enable\_alertmanager) | Creates Amazon Managed Service for Prometheus AlertManager for all workloads | `bool` | `false` | no | | <a name="input_enable_alertmanager"></a> [enable\_alertmanager](#input\_enable\_alertmanager) | Creates Amazon Managed Service for Prometheus AlertManager for all workloads | `bool` | `false` | no |
| <a name="input_enable_amazon_eks_adot"></a> [enable\_amazon\_eks\_adot](#input\_enable\_amazon\_eks\_adot) | Enables the ADOT Operator on the EKS Cluster | `bool` | `true` | no |
| <a name="input_enable_cert_manager"></a> [enable\_cert\_manager](#input\_enable\_cert\_manager) | Allow reusing an existing installation of cert-manager | `bool` | `true` | no |
| <a name="input_enable_managed_grafana"></a> [enable\_managed\_grafana](#input\_enable\_managed\_grafana) | Creates a new Amazon Managed Grafana Workspace | `bool` | `true` | no | | <a name="input_enable_managed_grafana"></a> [enable\_managed\_grafana](#input\_enable\_managed\_grafana) | Creates a new Amazon Managed Grafana Workspace | `bool` | `true` | no |
| <a name="input_enable_managed_prometheus"></a> [enable\_managed\_prometheus](#input\_enable\_managed\_prometheus) | Creates a new Amazon Managed Service for Prometheus Workspace | `bool` | `true` | no | | <a name="input_enable_managed_prometheus"></a> [enable\_managed\_prometheus](#input\_enable\_managed\_prometheus) | Creates a new Amazon Managed Service for Prometheus Workspace | `bool` | `true` | no |
| <a name="input_grafana_api_key"></a> [grafana\_api\_key](#input\_grafana\_api\_key) | Grafana API key for the Amazon Managed Grafana workspace | `string` | n/a | yes | | <a name="input_grafana_api_key"></a> [grafana\_api\_key](#input\_grafana\_api\_key) | Grafana API key for the Amazon Managed Grafana workspace | `string` | n/a | yes |
| <a name="input_irsa_iam_permissions_boundary"></a> [irsa\_iam\_permissions\_boundary](#input\_irsa\_iam\_permissions\_boundary) | IAM permissions boundary for IRSA roles | `string` | `null` | no |
| <a name="input_irsa_iam_role_path"></a> [irsa\_iam\_role\_path](#input\_irsa\_iam\_role\_path) | IAM role path for IRSA roles | `string` | `"/"` | no |
| <a name="input_managed_grafana_workspace_id"></a> [managed\_grafana\_workspace\_id](#input\_managed\_grafana\_workspace\_id) | Amazon Managed Grafana Workspace ID | `string` | `""` | no | | <a name="input_managed_grafana_workspace_id"></a> [managed\_grafana\_workspace\_id](#input\_managed\_grafana\_workspace\_id) | Amazon Managed Grafana Workspace ID | `string` | `""` | no |
| <a name="input_managed_prometheus_workspace_id"></a> [managed\_prometheus\_workspace\_id](#input\_managed\_prometheus\_workspace\_id) | Amazon Managed Service for Prometheus Workspace ID | `string` | `""` | no | | <a name="input_managed_prometheus_workspace_id"></a> [managed\_prometheus\_workspace\_id](#input\_managed\_prometheus\_workspace\_id) | Amazon Managed Service for Prometheus Workspace ID | `string` | `""` | no |
| <a name="input_managed_prometheus_workspace_region"></a> [managed\_prometheus\_workspace\_region](#input\_managed\_prometheus\_workspace\_region) | Region where Amazon Managed Service for Prometheus is deployed | `string` | `null` | no | | <a name="input_managed_prometheus_workspace_region"></a> [managed\_prometheus\_workspace\_region](#input\_managed\_prometheus\_workspace\_region) | Region where Amazon Managed Service for Prometheus is deployed | `string` | `null` | no |
@@ -183,9 +182,8 @@ If you are interested in contributing, see the [Contribution guide](https://gith
| Name | Description | | Name | Description |
|------|-------------| |------|-------------|
| <a name="output_aws_region"></a> [aws\_region](#output\_aws\_region) | AWS Region | | <a name="output_aws_region"></a> [aws\_region](#output\_aws\_region) | AWS Region |
| <a name="output_eks_cluster_id"></a> [eks\_cluster\_id](#output\_eks\_cluster\_id) | EKS Cluster Id |
| <a name="output_eks_cluster_version"></a> [eks\_cluster\_version](#output\_eks\_cluster\_version) | EKS Cluster version |
| <a name="output_grafana_dashboards_folder_id"></a> [grafana\_dashboards\_folder\_id](#output\_grafana\_dashboards\_folder\_id) | Grafana folder ID for automatic dashboards. Required by workload modules | | <a name="output_grafana_dashboards_folder_id"></a> [grafana\_dashboards\_folder\_id](#output\_grafana\_dashboards\_folder\_id) | Grafana folder ID for automatic dashboards. Required by workload modules |
| <a name="output_grafana_prometheus_datasource_test"></a> [grafana\_prometheus\_datasource\_test](#output\_grafana\_prometheus\_datasource\_test) | Grafana save & test URL for Amazon Managed Prometheus workspace |
| <a name="output_managed_grafana_workspace_endpoint"></a> [managed\_grafana\_workspace\_endpoint](#output\_managed\_grafana\_workspace\_endpoint) | Amazon Managed Grafana workspace endpoint | | <a name="output_managed_grafana_workspace_endpoint"></a> [managed\_grafana\_workspace\_endpoint](#output\_managed\_grafana\_workspace\_endpoint) | Amazon Managed Grafana workspace endpoint |
| <a name="output_managed_grafana_workspace_id"></a> [managed\_grafana\_workspace\_id](#output\_managed\_grafana\_workspace\_id) | Amazon Managed Grafana workspace ID | | <a name="output_managed_grafana_workspace_id"></a> [managed\_grafana\_workspace\_id](#output\_managed\_grafana\_workspace\_id) | Amazon Managed Grafana workspace ID |
| <a name="output_managed_prometheus_workspace_endpoint"></a> [managed\_prometheus\_workspace\_endpoint](#output\_managed\_prometheus\_workspace\_endpoint) | Amazon Managed Prometheus workspace endpoint | | <a name="output_managed_prometheus_workspace_endpoint"></a> [managed\_prometheus\_workspace\_endpoint](#output\_managed\_prometheus\_workspace\_endpoint) | Amazon Managed Prometheus workspace endpoint |
+16 -4
View File
@@ -31,6 +31,18 @@ you need to track changes as part of a Git repository or CI/CD pipeline.
!!! warning !!! warning
When using `tfvars` files, always be careful to not store and commit any secrets (keys, passwords, ...) When using `tfvars` files, always be careful to not store and commit any secrets (keys, passwords, ...)
## v2.x changes
v2.x [releases](https://github.com/aws-observability/terraform-aws-observability-accelerator/releases) introduce
couple of breaking changes compared to previous versions:
- `modules/workloads/infra` module moves to `modules/eks-monitoring`
- EKS configuration options moves from the base module to the `eks-monitoring` module
- EKS workload modules **java,nginx** merge into `eks-monitoring` as configuration options (patterns),
see [examples](https://github.com/aws-observability/terraform-aws-observability-accelerator/tree/main/examples)
- Examples have been updated to reflect these changes
## Base module ## Base module
The base module allows you to configure the AWS Observability services for your cluster and The base module allows you to configure the AWS Observability services for your cluster and
@@ -41,7 +53,7 @@ and ADOT Operator deployed for you and ready to receive your data.
The base module serve as an anchor to the workload modules and cannot run on its own. The base module serve as an anchor to the workload modules and cannot run on its own.
```hcl ```hcl
module "eks_observability_accelerator" { module "aws_observability_accelerator" {
# use release tags and check for the latest versions # use release tags and check for the latest versions
# https://github.com/aws-observability/terraform-aws-observability-accelerator/releases # https://github.com/aws-observability/terraform-aws-observability-accelerator/releases
source = "github.com/aws-observability/terraform-aws-observability-accelerator?ref=v1.6.1" source = "github.com/aws-observability/terraform-aws-observability-accelerator?ref=v1.6.1"
@@ -49,7 +61,7 @@ module "eks_observability_accelerator" {
aws_region = "eu-west-1" aws_region = "eu-west-1"
eks_cluster_id = "my-eks-cluster" eks_cluster_id = "my-eks-cluster"
# As Grafana shares a different lifecycle, it's best to use an existing workspace. # As Grafana shares a different lifecycle, we recommend using an existing workspace.
managed_grafana_workspace_id = var.managed_grafana_workspace_id managed_grafana_workspace_id = var.managed_grafana_workspace_id
grafana_api_key = var.grafana_api_key grafana_api_key = var.grafana_api_key
} }
@@ -58,7 +70,7 @@ module "eks_observability_accelerator" {
You can optionally reuse an existing Amazon Managed Service for Prometheus Workspace: You can optionally reuse an existing Amazon Managed Service for Prometheus Workspace:
```hcl ```hcl
module "eks_observability_accelerator" { module "aws_observability_accelerator" {
# use release tags and check for the latest versions # use release tags and check for the latest versions
# https://github.com/aws-observability/terraform-aws-observability-accelerator/releases # https://github.com/aws-observability/terraform-aws-observability-accelerator/releases
source = "github.com/aws-observability/terraform-aws-observability-accelerator?ref=v1.6.1" source = "github.com/aws-observability/terraform-aws-observability-accelerator?ref=v1.6.1"
@@ -83,7 +95,7 @@ View all the configuration options in the [module's documentation](https://githu
Workloads modules are focused Terraform modules provided in this repository. They essentially provide curated metrics collection, alerts and Grafana dashboards according to the use case. Most of those modules require the base module. Workloads modules are focused Terraform modules provided in this repository. They essentially provide curated metrics collection, alerts and Grafana dashboards according to the use case. Most of those modules require the base module.
You can check the full workload modules list and their documentation [here](https://github.com/aws-observability/terraform-aws-observability-accelerator/tree/main/modules/workloads). You can check the full workload modules list and their documentation [here](https://github.com/aws-observability/terraform-aws-observability-accelerator/tree/main/modules/).
All the modules come with end-to-end deployable examples. All the modules come with end-to-end deployable examples.
+2 -4
View File
@@ -15,11 +15,9 @@ terraform destroy
To remove resources from your Terraform state, run To remove resources from your Terraform state, run
```bash ```bash
# grafana workspace
terraform state rm "module.eks_observability_accelerator.module.managed_grafana[0].aws_grafana_workspace.this[0]"
# prometheus workspace # prometheus workspace
terraform state rm "module.eks_observability_accelerator.aws_prometheus_workspace.this[0]" terraform state rm "module.eks_observability_accelerator.aws_prometheus_workspace.this[0]"
``` ```
> **Note:** To view all the features proposed by this module, visit the [module documentation](https://github.com/aws-observability/terraform-aws-observability-accelerator/tree/main/modules/workloads/infra). !!! note
To view all the features proposed by this module, visit the [module documentation](https://github.com/aws-observability/terraform-aws-observability-accelerator/tree/main/modules/workloads/infra).
+15 -16
View File
@@ -2,7 +2,7 @@
This example demonstrates how to monitor your Amazon Elastic Kubernetes Service This example demonstrates how to monitor your Amazon Elastic Kubernetes Service
(Amazon EKS) cluster with the Observability Accelerator's EKS (Amazon EKS) cluster with the Observability Accelerator's EKS
[infrastructure module](https://github.com/aws-observability/terraform-aws-observability-accelerator/tree/main/modules/workloads/infra). [infrastructure module](https://github.com/aws-observability/terraform-aws-observability-accelerator/tree/feat/modules-composition/modules/eks-monitoring).
Monitoring Amazon Elastic Kubernetes Service (Amazon EKS) for metrics has two categories: Monitoring Amazon Elastic Kubernetes Service (Amazon EKS) for metrics has two categories:
the control plane and the Amazon EKS nodes (with Kubernetes objects). the control plane and the Amazon EKS nodes (with Kubernetes objects).
@@ -72,24 +72,17 @@ aws amp create-workspace --alias observability-accelerator --query '.workspaceId
### 5. Amazon Managed Grafana workspace ### 5. Amazon Managed Grafana workspace
To run this example you need an Amazon Managed Grafana workspace. If you have an existing workspace, edit and run: To run this example you need an Amazon Managed Grafana workspace. If you have an existing workspace, create an environment variable as described below.
To create a new workspace, visit our Amazon Managed Grafana [documentation](https://docs.aws.amazon.com/grafana/latest/userguide/getting-started-with-AMG.html).
Make sure to provide the workspace with Amazon Managed Service for Prometheus read permissions.
!!! note
For the URL `https://g-xyz.grafana-workspace.eu-central-1.amazonaws.com`, the workspace ID would be `g-xyz`
```bash ```bash
export TF_VAR_managed_grafana_workspace_id=g-xxx export TF_VAR_managed_grafana_workspace_id=g-xxx
``` ```
To create a new one, within this example's Terraform state (sharing the same lifecycle with all the
other resources created by Terraform):
- Edit main.tf and set `enable_managed_grafana = true`
- Run
```bash
terraform init
terraform apply -target "module.eks_observability_accelerator.module.managed_grafana[0].aws_grafana_workspace.this[0]"
export TF_VAR_managed_grafana_workspace_id=$(terraform output --raw managed_grafana_workspace_id)
```
### 6. Grafana API Key ### 6. Grafana API Key
Amazon Managed Grafana provides a control plane API for generating Grafana API keys. Amazon Managed Grafana provides a control plane API for generating Grafana API keys.
@@ -114,7 +107,13 @@ terraform apply
1. Prometheus datasource on Grafana 1. Prometheus datasource on Grafana
Open your Grafana workspace and under Configuration -> Data sources, you should see `aws-observability-accelerator`. Open and click `Save & test`. You should see a notification confirming that the Amazon Managed Service for Prometheus workspace is ready to be used on Grafana. Make sure to open the link in the output. After a successful deployment, this will open
the Prometheus datasource configuration on Grafana.
Click `Save & test` and you should see a notification confirming that the Amazon Managed Service for Prometheus workspace is ready to be used on Grafana.
```bash
terraform output grafana_prometheus_datasource_test
```
2. Grafana dashboards 2. Grafana dashboards
@@ -126,7 +125,7 @@ Open a specific dashboard and you should be able to view its visualization
<img width="2056" alt="cluster headlines" src="https://user-images.githubusercontent.com/10175027/199110753-9bc7a9b7-1b45-4598-89d3-32980154080e.png"> <img width="2056" alt="cluster headlines" src="https://user-images.githubusercontent.com/10175027/199110753-9bc7a9b7-1b45-4598-89d3-32980154080e.png">
2. Amazon Managed Service for Prometheus rules and alerts 3. Amazon Managed Service for Prometheus rules and alerts
Open the Amazon Managed Service for Prometheus console and view the details of your workspace. Under the `Rules management` tab, you should find new rules deployed. Open the Amazon Managed Service for Prometheus console and view the details of your workspace. Under the `Rules management` tab, you should find new rules deployed.
+120
View File
@@ -0,0 +1,120 @@
# Monitor Java/JMX applications running on Amazon EKS
!!! note
Since v2.x, Java based applications monitoring on EKS has been merged within
the [eks-monitoring module](https://github.com/aws-observability/terraform-aws-observability-accelerator/tree/feat/modules-composition/modules/eks-monitoring)
to allow visibility both on the cluster and the workloads, [#59](https://github.com/aws-observability/terraform-aws-observability-accelerator/issues/59).
In addition to EKS infrastructure monitoring, the current example provides
curated Grafana dashboards, Prometheus alerting and recording rules with multiple
configuration options for Java based workloads on EKS.
## Setup
### 1. Add Java metrics, dashboards and alerts
From the [previous example's](https://aws-observability.github.io/terraform-aws-observability-accelerator/eks/) configuration,
simply enable the Java pattern's flag.
```hcl
module "eks_monitoring" {
...
enable_java = true
}
```
You can further customize the Java pattern by providing `java_config` [options](https://github.com/aws-observability/terraform-aws-observability-accelerator/blob/feat/modules-composition/modules/eks-monitoring/README.md#input_java_config).
### 2. Grafana API key
Make sure to refresh your temporary Grafana API key
```bash
export TF_VAR_managed_grafana_workspace_id=g-xxx
export TF_VAR_grafana_api_key=`aws grafana create-workspace-api-key --key-name "observability-accelerator-$(date +%s)" --key-role ADMIN --seconds-to-live 1200 --workspace-id $TF_VAR_managed_grafana_workspace_id --query key --output text`
```
## Deploy
Simply run this command to deploy.
```bash
terraform apply
```
!!! note
To see the complete Java example, open the [example on the repository](https://github.com/aws-observability/terraform-aws-observability-accelerator/tree/main/examples/existing-cluster-java)
## Visualization
1. Grafana dashboards
Go to the Dashboards panel of your Grafana workspace. There will be a folder called `Observability Accelerator Dashboards`
<img width="832" alt="image" src="https://user-images.githubusercontent.com/97046295/194903648-57c55d30-6f90-4b03-9eb6-577aaba7dc22.png">
Open the "Java/JMX" dashboard to view its visualization
<img width="2560" alt="Grafana Java dashboard" src="https://user-images.githubusercontent.com/10175027/217821001-2119c81f-94bd-4811-8bbb-caaf1ae5a77a.png">
2. Amazon Managed Service for Prometheus rules and alerts
Open the Amazon Managed Service for Prometheus console and view the details of your workspace. Under the `Rules management` tab, you will find new rules deployed.
<img width="1314" alt="image" src="https://user-images.githubusercontent.com/97046295/194904104-09a28577-d149-478e-b0a1-dc21cb7effc1.png">
!!! note
To setup your alert receiver, with Amazon SNS, follow [this documentation](https://docs.aws.amazon.com/prometheus/latest/userguide/AMP-alertmanager-receiver.html)
## Deploy an example Java application
In this section we will reuse an example from the AWS OpenTelemetry collector [repository](https://github.com/aws-observability/aws-otel-collector/blob/main/docs/developers/container-insights-eks-jmx.md). For convenience, the steps can be found below.
1. Clone [this repository](https://github.com/aws-observability/aws-otel-test-framework) and navigate to the `sample-apps/jmx/` directory.
2. Authenticate to Amazon ECR
```sh
export AWS_ACCOUNT_ID=`aws sts get-caller-identity --query Account --output text`
export AWS_REGION={region}
aws ecr get-login-password --region $AWS_REGION | docker login --username AWS --password-stdin $AWS_ACCOUNT_ID.dkr.ecr.$AWS_REGION.amazonaws.com
```
3. Create an Amazon ECR repository
```sh
aws ecr create-repository --repository-name prometheus-sample-tomcat-jmx \
--image-scanning-configuration scanOnPush=true \
--region $AWS_REGION
```
4. Build Docker image and push to ECR.
```sh
docker build -t $AWS_ACCOUNT_ID.dkr.ecr.$AWS_REGION.amazonaws.com/prometheus-sample-tomcat-jmx:latest .
docker push $AWS_ACCOUNT_ID.dkr.ecr.$AWS_REGION.amazonaws.com/prometheus-sample-tomcat-jmx:latest
```
5. Install sample application
```sh
export SAMPLE_TRAFFIC_NAMESPACE=javajmx-sample
curl https://raw.githubusercontent.com/aws-observability/aws-otel-test-framework/terraform/sample-apps/jmx/examples/prometheus-metrics-sample.yaml > metrics-sample.yaml
sed -i "s/{{aws_account_id}}/$AWS_ACCOUNT_ID/g" metrics-sample.yaml
sed -i "s/{{region}}/$AWS_REGION/g" metrics-sample.yaml
sed -i "s/{{namespace}}/$SAMPLE_TRAFFIC_NAMESPACE/g" metrics-sample.yaml
kubectl apply -f metrics-sample.yaml
```
Verify that the sample application is running:
```sh
kubectl get pods -n $SAMPLE_TRAFFIC_NAMESPACE
NAME READY STATUS RESTARTS AGE
tomcat-bad-traffic-generator 1/1 Running 0 11s
tomcat-example-7958666589-2q755 0/1 ContainerCreating 0 11s
tomcat-traffic-generator 1/1 Running 0 11s
```
+120
View File
@@ -0,0 +1,120 @@
# Monitor Nginx applications running on Amazon EKS
!!! note
Since v2.x, NGINX based applications monitoring on EKS has been merged within
the [eks-monitoring module](https://github.com/aws-observability/terraform-aws-observability-accelerator/tree/feat/modules-composition/modules/eks-monitoring)
to allow visibility both on the cluster and the workloads, [#59](https://github.com/aws-observability/terraform-aws-observability-accelerator/issues/59).
In addition to EKS infrastructure monitoring, the current example provides
curated Grafana dashboards, Prometheus alerting and recording rules with multiple
configuration options for NGINX based workloads on EKS.
## Setup
### 1. Add NGINX metrics, dashboards and alerts
From the [EKS cluster monitoring example's](https://aws-observability.github.io/terraform-aws-observability-accelerator/eks/) configuration,
simply enable the NGINX pattern's flag.
```hcl
module "eks_monitoring" {
...
enable_nginx = true
}
```
You can further customize the NGINX pattern by providing `nginx_config` [options](https://github.com/aws-observability/terraform-aws-observability-accelerator/blob/feat/modules-composition/modules/eks-monitoring/README.md#input_nginx_config).
### 2. Grafana API key
Make sure to refresh your temporary Grafana API key
```bash
export TF_VAR_managed_grafana_workspace_id=g-xxx
export TF_VAR_grafana_api_key=`aws grafana create-workspace-api-key --key-name "observability-accelerator-$(date +%s)" --key-role ADMIN --seconds-to-live 1200 --workspace-id $TF_VAR_managed_grafana_workspace_id --query key --output text`
```
## Deploy
Simply run this command to deploy.
```bash
terraform apply
```
!!! note
To see the complete NGINX example, open the [example on the repository](https://github.com/aws-observability/terraform-aws-observability-accelerator/tree/main/examples/existing-cluster-nginx)
## Visualization
1. Grafana dashboards
Go to the Dashboards panel of your Grafana workspace. You will see a list of dashboards under the `Observability Accelerator Dashboards`
<img width="1208" alt="image" src="https://user-images.githubusercontent.com/97046295/190665211-60faef71-d83d-4d59-ac80-bf4309d8c082.png">
Open the NGINX dashboard and you will be able to view its visualization
<img width="1850" alt="image" src="https://user-images.githubusercontent.com/97046295/196226043-e49afeb9-7828-467f-9199-5707cdc69aa9.png">
2. Amazon Managed Service for Prometheus rules and alerts
Open the Amazon Managed Service for Prometheus console and view the details of your workspace. Under the `Rules management` tab, you will find new rules deployed.
<img width="1054" alt="image" src="https://user-images.githubusercontent.com/97046295/190665728-ae8bb709-ad93-4629-b845-85c158dd1925.png">
!!! note
To setup your alert receiver, with Amazon SNS, follow [this documentation](https://docs.aws.amazon.com/prometheus/latest/userguide/AMP-alertmanager-receiver.html)
## Deploy an example application to visualize metrics
In this section we will deploy sample application and extract metrics using AWS OpenTelemetry collector
### 1. Add the helm incubator repo:
```sh
helm repo add ingress-nginx https://kubernetes.github.io/ingress-nginx
```
### 2. Enter the following command to create a new namespace:
```sh
kubectl create namespace nginx-ingress-sample
```
### 3. Enter the following commands to install NGINX:
```sh
helm install my-nginx ingress-nginx/ingress-nginx \
--namespace nginx-ingress-sample \
--set controller.metrics.enabled=true \
--set-string controller.metrics.service.annotations."prometheus\.io/port"="10254" \
--set-string controller.metrics.service.annotations."prometheus\.io/scrape"="true"
```
### 4. Set an EXTERNAL-IP variable to the value of the EXTERNAL-IP column in the row of the NGINX ingress controller.
```sh
EXTERNAL_IP=your-nginx-controller-external-ip
```
### 5. Start some sample NGINX traffic by entering the following command.
```sh
SAMPLE_TRAFFIC_NAMESPACE=nginx-sample-traffic
curl https://raw.githubusercontent.com/aws-samples/amazon-cloudwatch-container-insights/master/k8s-deployment-manifest-templates/deployment-mode/service/cwagent-prometheus/sample_traffic/nginx-traffic/nginx-traffic-sample.yaml |
sed "s/{{external_ip}}/$EXTERNAL_IP/g" |
sed "s/{{namespace}}/$SAMPLE_TRAFFIC_NAMESPACE/g" |
kubectl apply -f -
```
### 6. Verify if the application is running
```sh
kubectl get pods -n nginx-ingress-sample
```
### 7. Visualize the Application's dashboard
Log back into your Managed Grafana Workspace and navigate to the dashboard side panel, click on `Observability Accelerator Dashboards` Folder and open the `NGINX` Dashboard.
+1 -1
View File
@@ -31,7 +31,7 @@ to be deployed in our packaged
We have supporting examples for quick setup such as: We have supporting examples for quick setup such as:
- Creating an empty Amazon EKS cluster and a VPC - Creating an empty Amazon EKS cluster and a VPC
- Creating and configure an Amazon Managed Grafana workspace with SSO - Creating and configure an Amazon Managed Grafana workspace with SSO (coming soon)
## Motivation ## Motivation
+9
View File
@@ -0,0 +1,9 @@
# Support & Feedback
AWS Observability Accelerator for Terraform is maintained by AWS Solution Architects.
It is not part of an AWS service and support is provided best-effort by the
AWS Observability Accelerator community.
To post feedback, submit feature ideas, or report bugs, please use the [issues](https://github.com/aws-observability/terraform-aws-observability-accelerator/issues) section of this GitHub repo.
If you are interested in contributing, see the [contribution guide](https://github.com/aws-observability/terraform-aws-observability-accelerator/blob/main/CONTRIBUTING.md).
-196
View File
@@ -1,196 +0,0 @@
# Monitor Java/JMX applications running on Amazon EKS
The current example deploys the [java workload module](https://github.com/aws-observability/terraform-aws-observability-accelerator/tree/main/modules/workloads/java),
to provide to an existing EKS cluster with an OpenTelemetry collector,
curated Grafana dashboards, Prometheus alerting and recording rules with multiple
configuration options on the cluster infrastructure.
## Prerequisites
!!! note
Make sure to complete the [prerequisites section](https://aws-observability.github.io/terraform-aws-observability-accelerator/concepts/#prerequisites) before proceeding.
## Setup
### 1. Download sources and initialize Terraform
```bash
git clone https://github.com/aws-observability/terraform-aws-observability-accelerator.git
cd examples/existing-cluster-java
terraform init
```
### 2. AWS Region
Specify the AWS Region where the resources will be deployed:
```bash
export TF_VAR_aws_region=xxx
```
### 3. Amazon EKS Cluster
To run this example, you need to provide your EKS cluster name. If you don't
have a cluster ready, visit [this example](https://aws-observability.github.io/terraform-aws-observability-accelerator/helpers/new-eks-cluster/)
first to create a new one.
Specify your cluster name:
```bash
export TF_VAR_eks_cluster_id=xxx
```
### 4. Amazon Managed Service for Prometheus workspace (optional)
By default, we create an Amazon Managed Service for Prometheus workspace for you.
However, if you have an existing workspace you want to reuse, edit and run:
```bash
export TF_VAR_managed_prometheus_workspace_id=ws-xxx
```
To create a workspace outside of Terraform's state, simply run:
```bash
aws amp create-workspace --alias observability-accelerator --query '.workspaceId' --output text
```
### 5. Amazon Managed Grafana workspace
To run this example you need an Amazon Managed Grafana workspace. If you have an existing workspace, edit and run:
```bash
export TF_VAR_managed_grafana_workspace_id=g-xxx
```
To create a new one, within this example's Terraform state (sharing the same lifecycle with all the other resources):
- Edit main.tf and set `enable_managed_grafana = true`
- Run
```bash
terraform init
terraform apply -target "module.eks_observability_accelerator.module.managed_grafana[0].aws_grafana_workspace.this[0]"
export TF_VAR_managed_grafana_workspace_id=$(terraform output --raw managed_grafana_workspace_id)
```
### 6. Grafana API Key
Amazon Managed Grafana provides a control plane API for generating Grafana API keys.
As a security best practice, we will provide to Terraform a short lived API key to
run the `apply` or `destroy` command.
Ensure you have necessary IAM permissions (`CreateWorkspaceApiKey, DeleteWorkspaceApiKey`)
```bash
export TF_VAR_grafana_api_key=`aws grafana create-workspace-api-key --key-name "observability-accelerator-$(date +%s)" --key-role ADMIN --seconds-to-live 1200 --workspace-id $TF_VAR_managed_grafana_workspace_id --query key --output text`
```
## Deploy
Simply run this command to deploy.
```bash
terraform apply
```
## Visualization
1. Prometheus datasource on Grafana
Open your Grafana workspace and under Configuration -> Data sources, you will see `aws-observability-accelerator`. Open and click `Save & test`. You will then see a notification confirming that the Amazon Managed Service for Prometheus workspace is ready to be used on Grafana.
2. Grafana dashboards
Go to the Dashboards panel of your Grafana workspace. There will be a folder called `Observability Accelerator Dashboards`
<img width="832" alt="image" src="https://user-images.githubusercontent.com/97046295/194903648-57c55d30-6f90-4b03-9eb6-577aaba7dc22.png">
Open the "Java/JMX" dashboard to view its visualization
<img width="2560" alt="Grafana Java dashboard" src="https://user-images.githubusercontent.com/10175027/217821001-2119c81f-94bd-4811-8bbb-caaf1ae5a77a.png">
2. Amazon Managed Service for Prometheus rules and alerts
Open the Amazon Managed Service for Prometheus console and view the details of your workspace. Under the `Rules management` tab, you will find new rules deployed.
<img width="1314" alt="image" src="https://user-images.githubusercontent.com/97046295/194904104-09a28577-d149-478e-b0a1-dc21cb7effc1.png">
!!! note
To setup your alert receiver, with Amazon SNS, follow [this documentation](https://docs.aws.amazon.com/prometheus/latest/userguide/AMP-alertmanager-receiver.html)
## Deploy an Example Java Application
In this section we will reuse an example from the AWS OpenTelemetry collector [repository](https://github.com/aws-observability/aws-otel-collector/blob/main/docs/developers/container-insights-eks-jmx.md). For convenience, the steps can be found below.
1. Clone [this repository](https://github.com/aws-observability/aws-otel-test-framework) and navigate to the `sample-apps/jmx/` directory.
2. Authenticate to Amazon ECR
```sh
export AWS_ACCOUNT_ID=`aws sts get-caller-identity --query Account --output text`
export AWS_REGION={region}
aws ecr get-login-password --region $AWS_REGION | docker login --username AWS --password-stdin $AWS_ACCOUNT_ID.dkr.ecr.$AWS_REGION.amazonaws.com
```
3. Create an Amazon ECR repository
```sh
aws ecr create-repository --repository-name prometheus-sample-tomcat-jmx \
--image-scanning-configuration scanOnPush=true \
--region $AWS_REGION
```
4. Build Docker image and push to ECR.
```sh
docker build -t $AWS_ACCOUNT_ID.dkr.ecr.$AWS_REGION.amazonaws.com/prometheus-sample-tomcat-jmx:latest .
docker push $AWS_ACCOUNT_ID.dkr.ecr.$AWS_REGION.amazonaws.com/prometheus-sample-tomcat-jmx:latest
```
5. Install sample application
```sh
export SAMPLE_TRAFFIC_NAMESPACE=javajmx-sample
curl https://raw.githubusercontent.com/aws-observability/aws-otel-test-framework/terraform/sample-apps/jmx/examples/prometheus-metrics-sample.yaml > metrics-sample.yaml
sed -i "s/{{aws_account_id}}/$AWS_ACCOUNT_ID/g" metrics-sample.yaml
sed -i "s/{{region}}/$AWS_REGION/g" metrics-sample.yaml
sed -i "s/{{namespace}}/$SAMPLE_TRAFFIC_NAMESPACE/g" metrics-sample.yaml
kubectl apply -f metrics-sample.yaml
```
Verify that the sample application is running:
```sh
kubectl get pods -n $SAMPLE_TRAFFIC_NAMESPACE
NAME READY STATUS RESTARTS AGE
tomcat-bad-traffic-generator 1/1 Running 0 11s
tomcat-example-7958666589-2q755 0/1 ContainerCreating 0 11s
tomcat-traffic-generator 1/1 Running 0 11s
```
## Destroy resources
If you leave this stack running, you will continue to incur charges. To remove all resources
created by Terraform, [refresh your Grafana API key](#6-grafana-api-key) and run the command below.
!!! warning
Be careful, this command will removing everything created by Terraform. If you wish
to keep your Amazon Managed Grafana or Amazon Managed Service for Prometheus workspaces. Remove them
from your terraform state before running the destroy command.
```bash
terraform destroy
```
To remove resources from your Terraform state, run
```bash
# grafana workspace
terraform state rm "module.eks_observability_accelerator.module.managed_grafana[0].aws_grafana_workspace.this[0]"
# prometheus workspace
terraform state rm "module.eks_observability_accelerator.aws_prometheus_workspace.this[0]"
```
-198
View File
@@ -1,198 +0,0 @@
# Monitor Nginx applications running on Amazon EKS
The current example deploys the [nginx workload module](https://github.com/aws-observability/terraform-aws-observability-accelerator/tree/main/modules/workloads/nginx),
to provide an existing EKS cluster with an OpenTelemetry collector,
curated Grafana dashboards, Prometheus alerting and recording rules with multiple
configuration options on the cluster infrastructure.
## Prerequisites
!!! note
Make sure to complete the [prerequisites section](https://aws-observability.github.io/terraform-aws-observability-accelerator/concepts/#prerequisites) before proceeding.
## Setup
### 1. Download sources and initialize Terraform
```bash
git clone https://github.com/aws-observability/terraform-aws-observability-accelerator.git
cd examples/existing-cluster-nginx
terraform init
```
### 2. AWS Region
Specify the AWS Region where the resources will be deployed:
```bash
export TF_VAR_aws_region=xxx
```
### 3. Amazon EKS Cluster
To run this example, you need to provide your EKS cluster name. If you don't
have a cluster ready, visit [this example](https://aws-observability.github.io/terraform-aws-observability-accelerator/helpers/new-eks-cluster/)
first to create a new one.
Specify your cluster name:
```bash
export TF_VAR_eks_cluster_id=xxx
```
### 4. Amazon Managed Service for Prometheus workspace (optional)
By default, we create an Amazon Managed Service for Prometheus workspace for you.
However, if you have an existing workspace you want to reuse, edit and run:
```bash
export TF_VAR_managed_prometheus_workspace_id=ws-xxx
```
To create a workspace outside of Terraform's state, simply run:
```bash
aws amp create-workspace --alias observability-accelerator --query '.workspaceId' --output text
```
### 5. Amazon Managed Grafana workspace
To run this example you need an Amazon Managed Grafana workspace. If you have an existing workspace, edit and run:
```bash
export TF_VAR_managed_grafana_workspace_id=g-xxx
```
To create a new one, within this example's Terraform state (sharing the same lifecycle with all the other resources):
- Edit main.tf and set `enable_managed_grafana = true`
- Run
```bash
terraform init
terraform apply -target "module.eks_observability_accelerator.module.managed_grafana[0].aws_grafana_workspace.this[0]"
export TF_VAR_managed_grafana_workspace_id=$(terraform output --raw managed_grafana_workspace_id)
```
### 6. Grafana API Key
Amazon Managed Grafana provides a control plane API for generating Grafana API keys.
As a security best practice, we will provide to Terraform a short lived API key to
run the `apply` or `destroy` command.
Ensure you have necessary IAM permissions (`CreateWorkspaceApiKey, DeleteWorkspaceApiKey`)
```bash
export TF_VAR_grafana_api_key=`aws grafana create-workspace-api-key --key-name "observability-accelerator-$(date +%s)" --key-role ADMIN --seconds-to-live 1200 --workspace-id $TF_VAR_managed_grafana_workspace_id --query key --output text`
```
## Deploy
Simply run this command to deploy.
```bash
terraform apply
```
## Visualization
### 1. Prometheus datasource on Grafana
Open your Grafana workspace and under Configuration -> Data sources, you will see `aws-observability-accelerator`. Open and click `Save & test`. You will see a notification confirming that the Amazon Managed Service for Prometheus workspace is ready to be used on Grafana.
### 2. Grafana dashboards
Go to the Dashboards panel of your Grafana workspace. You will see a list of dashboards under the `Observability Accelerator Dashboards`
<img width="1208" alt="image" src="https://user-images.githubusercontent.com/97046295/190665211-60faef71-d83d-4d59-ac80-bf4309d8c082.png">
Open the NGINX dashboard and you will be able to view its visualization
<img width="1850" alt="image" src="https://user-images.githubusercontent.com/97046295/196226043-e49afeb9-7828-467f-9199-5707cdc69aa9.png">
### 3. Amazon Managed Service for Prometheus rules and alerts
Open the Amazon Managed Service for Prometheus console and view the details of your workspace. Under the `Rules management` tab, you will find new rules deployed.
<img width="1054" alt="image" src="https://user-images.githubusercontent.com/97046295/190665728-ae8bb709-ad93-4629-b845-85c158dd1925.png">
!!! note
To setup your alert receiver, with Amazon SNS, follow [this documentation](https://docs.aws.amazon.com/prometheus/latest/userguide/AMP-alertmanager-receiver.html)
## Deploy an Example Application to Visualize
In this section we will deploy sample application and extract metrics using AWS OpenTelemetry collector
### 1. Add the helm incubator repo:
```sh
helm repo add ingress-nginx https://kubernetes.github.io/ingress-nginx
```
### 2. Enter the following command to create a new namespace:
```sh
kubectl create namespace nginx-ingress-sample
```
### 3. Enter the following commands to install NGINX:
```sh
helm install my-nginx ingress-nginx/ingress-nginx \
--namespace nginx-ingress-sample \
--set controller.metrics.enabled=true \
--set-string controller.metrics.service.annotations."prometheus\.io/port"="10254" \
--set-string controller.metrics.service.annotations."prometheus\.io/scrape"="true"
```
### 4. Set an EXTERNAL-IP variable to the value of the EXTERNAL-IP column in the row of the NGINX ingress controller.
```sh
EXTERNAL_IP=your-nginx-controller-external-ip
```
### 5. Start some sample NGINX traffic by entering the following command.
```sh
SAMPLE_TRAFFIC_NAMESPACE=nginx-sample-traffic
curl https://raw.githubusercontent.com/aws-samples/amazon-cloudwatch-container-insights/master/k8s-deployment-manifest-templates/deployment-mode/service/cwagent-prometheus/sample_traffic/nginx-traffic/nginx-traffic-sample.yaml |
sed "s/{{external_ip}}/$EXTERNAL_IP/g" |
sed "s/{{namespace}}/$SAMPLE_TRAFFIC_NAMESPACE/g" |
kubectl apply -f -
```
### 6. Verify if the application is running
```sh
kubectl get pods -n nginx-ingress-sample
```
### 7. Visualize the Application's dashboard
Log back into your Managed Grafana Workspace and navigate to the dashboard side panel, click on `Observability Accelerator Dashboards` Folder and open the `NGINX` Dashboard.
## Destroy resources
If you leave this stack running, you will continue to incur charges. To remove all resources
created by Terraform, [refresh your Grafana API key](#6-grafana-api-key) and run the command below.
!!! warning
Be careful, this command will removing everything created by Terraform. If you wish
to keep your Amazon Managed Grafana or Amazon Managed Service for Prometheus workspaces. Remove them
from your terraform state before running the destroy command.
```bash
terraform destroy
```
To remove resources from your Terraform state, run
```bash
# grafana workspace
terraform state rm "module.eks_observability_accelerator.module.managed_grafana[0].aws_grafana_workspace.this[0]"
# prometheus workspace
terraform state rm "module.eks_observability_accelerator.aws_prometheus_workspace.this[0]"
```
+41 -38
View File
@@ -1,15 +1,16 @@
# Existing Cluster with the AWS Observability accelerator base module and Java monitoring # Monitor Java applications running on Amazon EKS
This example demonstrates how to use the AWS Observability Accelerator Terraform This example demonstrates how to use the AWS Observability Accelerator Terraform
modules with Java monitoring enabled. modules to monitor EKS infrastructure and Java based workloads.
The current example deploys the [AWS Distro for OpenTelemetry Operator](https://docs.aws.amazon.com/eks/latest/userguide/opentelemetry.html) for Amazon EKS with its requirements and make use of existing The current example deploys the [AWS Distro for OpenTelemetry Operator](https://docs.aws.amazon.com/eks/latest/userguide/opentelemetry.html)
Amazon Managed Service for Prometheus and Amazon Managed Grafana workspaces. for Amazon EKS with its requirements and make use of an existing Amazon Managed Grafana workspace.
It creates a new Amazon Managed Service for Prometheus workspace unless provided with an existing one to reuse.
It is based on the `java module`, one of our [workloads modules](../../modules/workloads/) Since v2.x releases, it uses the `EKS monitoring` [module](../../modules/eks-monitoring/)
to provide an existing EKS cluster with an OpenTelemetry collector, to provide an existing EKS cluster with an OpenTelemetry collector,
curated Grafana dashboards, Prometheus alerting and recording rules with multiple curated Grafana dashboards, Prometheus alerting and recording rules with multiple
configuration options on the cluster infrastructure. configuration options on the cluster infrastructure.
You will gain both visibility on the cluster and Java based applications.
## Prerequisites ## Prerequisites
@@ -39,11 +40,7 @@ cd examples/existing-cluster-java
terraform init terraform init
``` ```
3. AWS Region 3. Amazon EKS Cluster
Specify the AWS Region where the resources will be deployed. Edit the `terraform.tfvars` file and modify `aws_region="..."`. You can also use environement variables `export TF_VAR_aws_region=xxx`.
4. Amazon EKS Cluster
To run this example, you need to provide your EKS cluster name. To run this example, you need to provide your EKS cluster name.
If you don't have a cluster ready, visit [this example](../eks-cluster-with-vpc) If you don't have a cluster ready, visit [this example](../eks-cluster-with-vpc)
@@ -51,18 +48,15 @@ first to create a new one.
Add your cluster name for `eks_cluster_id="..."` to the `terraform.tfvars` or use an environment variable `export TF_VAR_eks_cluster_id=xxx`. Add your cluster name for `eks_cluster_id="..."` to the `terraform.tfvars` or use an environment variable `export TF_VAR_eks_cluster_id=xxx`.
5. Amazon Managed Service for Prometheus workspace (optional) 4. Amazon Managed Grafana workspace
If you have an existing workspace, add `managed_prometheus_workspace_id=ws-xxx` To run this example you need an Amazon Managed Grafana workspace. If you have an existing workspace, create an environment variable `export TF_VAR_managed_grafana_workspace_id=g-xxx`.
or use an environment variable `export TF_VAR_managed_prometheus_workspace_id=ws-xxx`. To create a new one, visit our Amazon Managed Grafana [documentation](https://docs.aws.amazon.com/grafana/latest/userguide/getting-started-with-AMG.html).
Make sure to provide the workspace with Amazon Managed Service for Prometheus read permissions.
If you don't specify anything a new workspace will be created for you. > In the URL `https://g-xyz.grafana-workspace.eu-central-1.amazonaws.com`, the workspace ID would be `g-xyz`
6. Amazon Managed Grafana workspace 5. <a name="apikey"></a> Grafana API Key
If you have an existing workspace, create an environment variable `export TF_VAR_managed_grafana_workspace_id=g-xxx`.
7. <a name="apikey"></a> Grafana API Key
Amazon Managed Service for Grafana provides a control plane API for generating Grafana API keys. We will provide to Terraform Amazon Managed Service for Grafana provides a control plane API for generating Grafana API keys. We will provide to Terraform
a short lived API key to run the `apply` or `destroy` command. a short lived API key to run the `apply` or `destroy` command.
@@ -84,11 +78,31 @@ or if you had only setup environment variables, run
terraform apply terraform apply
``` ```
## Additional configuration
For the purpose of the example, we have provided default values for some of the variables.
1. AWS Region
Specify the AWS Region where the resources will be deployed. Edit the `terraform.tfvars` file and modify `aws_region="..."`. You can also use environement variables `export TF_VAR_aws_region=xxx`.
2. Amazon Managed Service for Prometheus workspace
If you have an existing workspace, add `managed_prometheus_workspace_id=ws-xxx`
or use an environment variable `export TF_VAR_managed_prometheus_workspace_id=ws-xxx`.
## Visualization ## Visualization
1. Prometheus datasource on Grafana 1. Prometheus datasource on Grafana
Open your Grafana workspace and under Configuration -> Data sources, you will see `aws-observability-accelerator`. Open and click `Save & test`. You will then see a notification confirming that the Amazon Managed Service for Prometheus workspace is ready to be used on Grafana. Make sure to open the link in the output. After a successful deployment, this will open
the Prometheus datasource configuration on Grafana.
Click `Save & test` and you should see a notification confirming that the Amazon Managed Service for Prometheus workspace is ready to be used on Grafana.
```bash
terraform output grafana_prometheus_datasource_test
```
2. Grafana dashboards 2. Grafana dashboards
@@ -109,7 +123,7 @@ Open the Amazon Managed Service for Prometheus console and view the details of y
To setup your alert receiver, with Amazon SNS, follow [this documentation](https://docs.aws.amazon.com/prometheus/latest/userguide/AMP-alertmanager-receiver.html) To setup your alert receiver, with Amazon SNS, follow [this documentation](https://docs.aws.amazon.com/prometheus/latest/userguide/AMP-alertmanager-receiver.html)
## Deploy an Example Java Application ## Deploy an example Java application
In this section we will reuse an example from the AWS OpenTelemetry collector [repository](https://github.com/aws-observability/aws-otel-collector/blob/main/docs/developers/container-insights-eks-jmx.md). For convenience, the steps can be found below. In this section we will reuse an example from the AWS OpenTelemetry collector [repository](https://github.com/aws-observability/aws-otel-collector/blob/main/docs/developers/container-insights-eks-jmx.md). For convenience, the steps can be found below.
@@ -160,18 +174,6 @@ tomcat-example-7958666589-2q755 0/1 ContainerCreating 0 11s
tomcat-traffic-generator 1/1 Running 0 11s tomcat-traffic-generator 1/1 Running 0 11s
``` ```
## Advanced configuration
1. Cross-region Amazon Managed Prometheus workspace
If your existing Amazon Managed Prometheus workspace is in another AWS Region,
add this `managed_prometheus_region=xxx` and `managed_prometheus_workspace_id=ws-xxx`.
2. Cross-region Amazon Managed Grafana workspace
If your existing Amazon Managed Prometheus workspace is in another AWS Region,
add this `managed_prometheus_region=xxx` and `managed_prometheus_workspace_id=ws-xxx`.
## Destroy resources ## Destroy resources
If you leave this stack running, you will continue to incur charges. To remove all resources If you leave this stack running, you will continue to incur charges. To remove all resources
@@ -204,8 +206,8 @@ terraform destroy -var-file=terraform.tfvars
| Name | Source | Version | | Name | Source | Version |
|------|--------|---------| |------|--------|---------|
| <a name="module_eks_observability_accelerator"></a> [eks\_observability\_accelerator](#module\_eks\_observability\_accelerator) | ../../ | n/a | | <a name="module_aws_observability_accelerator"></a> [aws\_observability\_accelerator](#module\_aws\_observability\_accelerator) | ../../ | n/a |
| <a name="module_workloads_java"></a> [workloads\_java](#module\_workloads\_java) | ../../modules/workloads/java | n/a | | <a name="module_eks_monitoring"></a> [eks\_monitoring](#module\_eks\_monitoring) | ../../modules/eks-monitoring | n/a |
## Resources ## Resources
@@ -220,8 +222,8 @@ terraform destroy -var-file=terraform.tfvars
|------|-------------|------|---------|:--------:| |------|-------------|------|---------|:--------:|
| <a name="input_aws_region"></a> [aws\_region](#input\_aws\_region) | AWS Region | `string` | n/a | yes | | <a name="input_aws_region"></a> [aws\_region](#input\_aws\_region) | AWS Region | `string` | n/a | yes |
| <a name="input_eks_cluster_id"></a> [eks\_cluster\_id](#input\_eks\_cluster\_id) | Name of the EKS cluster | `string` | n/a | yes | | <a name="input_eks_cluster_id"></a> [eks\_cluster\_id](#input\_eks\_cluster\_id) | Name of the EKS cluster | `string` | n/a | yes |
| <a name="input_grafana_api_key"></a> [grafana\_api\_key](#input\_grafana\_api\_key) | API key for authorizing the Grafana provider to make changes to Amazon Managed Grafana | `string` | `""` | no | | <a name="input_grafana_api_key"></a> [grafana\_api\_key](#input\_grafana\_api\_key) | API key for authorizing the Grafana provider to make changes to Amazon Managed Grafana | `string` | n/a | yes |
| <a name="input_managed_grafana_workspace_id"></a> [managed\_grafana\_workspace\_id](#input\_managed\_grafana\_workspace\_id) | Amazon Managed Grafana Workspace ID | `string` | `""` | no | | <a name="input_managed_grafana_workspace_id"></a> [managed\_grafana\_workspace\_id](#input\_managed\_grafana\_workspace\_id) | Amazon Managed Grafana Workspace ID | `string` | n/a | yes |
| <a name="input_managed_prometheus_workspace_id"></a> [managed\_prometheus\_workspace\_id](#input\_managed\_prometheus\_workspace\_id) | Amazon Managed Service for Prometheus Workspace ID | `string` | `""` | no | | <a name="input_managed_prometheus_workspace_id"></a> [managed\_prometheus\_workspace\_id](#input\_managed\_prometheus\_workspace\_id) | Amazon Managed Service for Prometheus Workspace ID | `string` | `""` | no |
## Outputs ## Outputs
@@ -232,6 +234,7 @@ terraform destroy -var-file=terraform.tfvars
| <a name="output_eks_cluster_id"></a> [eks\_cluster\_id](#output\_eks\_cluster\_id) | EKS Cluster Id | | <a name="output_eks_cluster_id"></a> [eks\_cluster\_id](#output\_eks\_cluster\_id) | EKS Cluster Id |
| <a name="output_eks_cluster_version"></a> [eks\_cluster\_version](#output\_eks\_cluster\_version) | EKS Cluster version | | <a name="output_eks_cluster_version"></a> [eks\_cluster\_version](#output\_eks\_cluster\_version) | EKS Cluster version |
| <a name="output_grafana_dashboard_urls"></a> [grafana\_dashboard\_urls](#output\_grafana\_dashboard\_urls) | URLs for dashboards created | | <a name="output_grafana_dashboard_urls"></a> [grafana\_dashboard\_urls](#output\_grafana\_dashboard\_urls) | URLs for dashboards created |
| <a name="output_grafana_prometheus_datasource_test"></a> [grafana\_prometheus\_datasource\_test](#output\_grafana\_prometheus\_datasource\_test) | Grafana save & test URL for Amazon Managed Prometheus workspace |
| <a name="output_managed_prometheus_workspace_endpoint"></a> [managed\_prometheus\_workspace\_endpoint](#output\_managed\_prometheus\_workspace\_endpoint) | Amazon Managed Prometheus workspace endpoint | | <a name="output_managed_prometheus_workspace_endpoint"></a> [managed\_prometheus\_workspace\_endpoint](#output\_managed\_prometheus\_workspace\_endpoint) | Amazon Managed Prometheus workspace endpoint |
| <a name="output_managed_prometheus_workspace_id"></a> [managed\_prometheus\_workspace\_id](#output\_managed\_prometheus\_workspace\_id) | Amazon Managed Prometheus workspace ID | | <a name="output_managed_prometheus_workspace_id"></a> [managed\_prometheus\_workspace\_id](#output\_managed\_prometheus\_workspace\_id) | Amazon Managed Prometheus workspace ID |
<!-- END OF PRE-COMMIT-TERRAFORM DOCS HOOK --> <!-- END OF PRE-COMMIT-TERRAFORM DOCS HOOK -->
+17 -24
View File
@@ -34,28 +34,17 @@ locals {
} }
# deploys the base module # deploys the base module
module "eks_observability_accelerator" { module "aws_observability_accelerator" {
# source = "aws-observability/terrarom-aws-observability-accelerator"
source = "../../" source = "../../"
# source = "github.com/aws-observability/terraform-aws-observability-accelerator?ref=v2.0.0"
aws_region = var.aws_region aws_region = var.aws_region
eks_cluster_id = var.eks_cluster_id
# deploys AWS Distro for OpenTelemetry operator into the cluster
enable_amazon_eks_adot = true
# reusing existing certificate manager? defaults to true
enable_cert_manager = true
# creates a new Amazon Managed Prometheus workspace, defaults to true # creates a new Amazon Managed Prometheus workspace, defaults to true
enable_managed_prometheus = local.create_new_workspace enable_managed_prometheus = local.create_new_workspace
# reusing existing Amazon Managed Prometheus if specified # reusing existing Amazon Managed Prometheus if specified
managed_prometheus_workspace_id = var.managed_prometheus_workspace_id managed_prometheus_workspace_id = var.managed_prometheus_workspace_id
managed_prometheus_workspace_region = null # defaults to the current region, useful for cross region scenarios (same account)
# sets up the Amazon Managed Prometheus alert manager at the workspace level
enable_alertmanager = true
# reusing existing Amazon Managed Grafana workspace # reusing existing Amazon Managed Grafana workspace
enable_managed_grafana = false enable_managed_grafana = false
@@ -70,20 +59,24 @@ module "eks_observability_accelerator" {
# any provider blocks. # any provider blocks.
# This allows forcing dependency between base and workloads module # This allows forcing dependency between base and workloads module
provider "grafana" { provider "grafana" {
url = module.eks_observability_accelerator.managed_grafana_workspace_endpoint url = module.aws_observability_accelerator.managed_grafana_workspace_endpoint
auth = var.grafana_api_key auth = var.grafana_api_key
} }
module "workloads_java" { module "eks_monitoring" {
source = "../../modules/workloads/java" source = "../../modules/eks-monitoring"
# source = "github.com/aws-observability/terraform-aws-observability-accelerator//modules/eks-monitoring?ref=v2.0.0"
eks_cluster_id = module.eks_observability_accelerator.eks_cluster_id # enable java metrics collection, dashboards and alerts rules creation
enable_java = true
dashboards_folder_id = module.eks_observability_accelerator.grafana_dashboards_folder_id eks_cluster_id = var.eks_cluster_id
managed_prometheus_workspace_id = module.eks_observability_accelerator.managed_prometheus_workspace_id
managed_prometheus_workspace_endpoint = module.eks_observability_accelerator.managed_prometheus_workspace_endpoint dashboards_folder_id = module.aws_observability_accelerator.grafana_dashboards_folder_id
managed_prometheus_workspace_region = module.eks_observability_accelerator.managed_prometheus_workspace_region managed_prometheus_workspace_id = module.aws_observability_accelerator.managed_prometheus_workspace_id
managed_prometheus_workspace_endpoint = module.aws_observability_accelerator.managed_prometheus_workspace_endpoint
managed_prometheus_workspace_region = module.aws_observability_accelerator.managed_prometheus_workspace_region
# optional, defaults to 60s interval and 15s timeout # optional, defaults to 60s interval and 15s timeout
prometheus_config = { prometheus_config = {
@@ -95,6 +88,6 @@ module "workloads_java" {
tags = local.tags tags = local.tags
depends_on = [ depends_on = [
module.eks_observability_accelerator module.aws_observability_accelerator
] ]
} }
+19 -14
View File
@@ -1,29 +1,34 @@
output "eks_cluster_id" {
description = "EKS Cluster Id"
value = module.eks_observability_accelerator.eks_cluster_id
}
output "aws_region" { output "aws_region" {
description = "AWS Region" description = "AWS Region"
value = module.eks_observability_accelerator.aws_region value = module.aws_observability_accelerator.aws_region
}
output "eks_cluster_version" {
description = "EKS Cluster version"
value = module.eks_observability_accelerator.eks_cluster_version
} }
output "managed_prometheus_workspace_endpoint" { output "managed_prometheus_workspace_endpoint" {
description = "Amazon Managed Prometheus workspace endpoint" description = "Amazon Managed Prometheus workspace endpoint"
value = module.eks_observability_accelerator.managed_prometheus_workspace_endpoint value = module.aws_observability_accelerator.managed_prometheus_workspace_endpoint
} }
output "managed_prometheus_workspace_id" { output "managed_prometheus_workspace_id" {
description = "Amazon Managed Prometheus workspace ID" description = "Amazon Managed Prometheus workspace ID"
value = module.eks_observability_accelerator.managed_prometheus_workspace_id value = module.aws_observability_accelerator.managed_prometheus_workspace_id
} }
output "grafana_dashboard_urls" { output "grafana_dashboard_urls" {
description = "URLs for dashboards created" description = "URLs for dashboards created"
value = module.workloads_java.grafana_dashboard_urls value = module.eks_monitoring.grafana_dashboard_urls
}
output "grafana_prometheus_datasource_test" {
description = "Grafana save & test URL for Amazon Managed Prometheus workspace"
value = module.aws_observability_accelerator.grafana_prometheus_datasource_test
}
output "eks_cluster_version" {
description = "EKS Cluster version"
value = module.eks_monitoring.eks_cluster_version
}
output "eks_cluster_id" {
description = "EKS Cluster Id"
value = module.eks_monitoring.eks_cluster_id
} }
@@ -14,11 +14,9 @@ variable "managed_prometheus_workspace_id" {
variable "managed_grafana_workspace_id" { variable "managed_grafana_workspace_id" {
description = "Amazon Managed Grafana Workspace ID" description = "Amazon Managed Grafana Workspace ID"
type = string type = string
default = ""
} }
variable "grafana_api_key" { variable "grafana_api_key" {
description = "API key for authorizing the Grafana provider to make changes to Amazon Managed Grafana" description = "API key for authorizing the Grafana provider to make changes to Amazon Managed Grafana"
type = string type = string
default = ""
sensitive = true sensitive = true
} }
+48 -38
View File
@@ -1,16 +1,16 @@
# Existing Cluster with the AWS Observability accelerator base module and Nginx monitoring # Monitor NGINX applications running on Amazon EKS
This example demonstrates how to use the AWS Observability Accelerator Terraform This example demonstrates how to use the AWS Observability Accelerator Terraform
modules with Nginx monitoring enabled. modules to monitor EKS infrastructure and NGINX workloads.
The current example deploys the [AWS Distro for OpenTelemetry Operator](https://docs.aws.amazon.com/eks/latest/userguide/opentelemetry.html) for Amazon EKS with its requirements and make use of existing The current example deploys the [AWS Distro for OpenTelemetry Operator](https://docs.aws.amazon.com/eks/latest/userguide/opentelemetry.html)
Amazon Managed Service for Prometheus and Amazon Managed Grafana workspaces. for Amazon EKS with its requirements and make use of an existing Amazon Managed Grafana workspace.
It creates a new Amazon Managed Service for Prometheus workspace unless provided with an existing one to reuse.
It is based on the `nginx module`, one of our [workload modules](../../modules/workloads/) Since v2.x releases, it uses the `EKS monitoring` [module](../../modules/eks-monitoring/)
to provide an existing EKS cluster with an OpenTelemetry collector, to provide an existing EKS cluster with an OpenTelemetry collector,
curated Grafana dashboards, Prometheus alerting and recording rules with multiple curated Grafana dashboards, Prometheus alerting and recording rules with multiple
configuration options on the cluster infrastructure. configuration options on the cluster infrastructure.
You will gain both visibility on the cluster and NGINX based applications.
## Prerequisites ## Prerequisites
@@ -39,11 +39,7 @@ cd examples/existing-cluster-nginx
terraform init terraform init
``` ```
3. AWS Region 3. Amazon EKS Cluster
Specify the AWS Region where the resources will be deployed. Edit the `terraform.tfvars` file and modify `aws_region="..."`. You can also use environement variables `export TF_VAR_aws_region=xxx`.
4. Amazon EKS Cluster
To run this example, you need to provide your EKS cluster name. To run this example, you need to provide your EKS cluster name.
If you don't have a cluster ready, visit [this example](https://github.com/aws-ia/terraform-aws-eks-blueprints/tree/v4.13.1/examples/eks-cluster-with-new-vpc) If you don't have a cluster ready, visit [this example](https://github.com/aws-ia/terraform-aws-eks-blueprints/tree/v4.13.1/examples/eks-cluster-with-new-vpc)
@@ -51,30 +47,23 @@ first to create a new one.
Add your cluster name for `eks_cluster_id="..."` to the `terraform.tfvars` or use an environment variable `export TF_VAR_eks_cluster_id=xxx`. Add your cluster name for `eks_cluster_id="..."` to the `terraform.tfvars` or use an environment variable `export TF_VAR_eks_cluster_id=xxx`.
5. Amazon Managed Service for Prometheus workspace (optional) 4. Amazon Managed Grafana workspace
If you have an existing workspace, add `managed_prometheus_workspace_id=ws-xxx` To run this example you need an Amazon Managed Grafana workspace. If you have an existing workspace, create an environment variable `export TF_VAR_managed_grafana_workspace_id=g-xxx`.
or use an environment variable `export TF_VAR_managed_prometheus_workspace_id=ws-xxx`. To create a new one, visit our Amazon Managed Grafana [documentation](https://docs.aws.amazon.com/grafana/latest/userguide/getting-started-with-AMG.html).
Make sure to provide the workspace with Amazon Managed Service for Prometheus read permissions.
If you don't specify anything a new workspace will be created for you. > In the URL `https://g-xyz.grafana-workspace.eu-central-1.amazonaws.com`, the workspace ID would be `g-xyz`
6. Amazon Managed Grafana workspace 5. <a name="apikey"></a> Grafana API Key
If you have an existing workspace, add `managed_grafana_workspace_id=g-xxx` Amazon Managed Service for Grafana provides a control plane API for generating Grafana API keys. We will provide to Terraform
or use an environment variable `export TF_VAR_managed_grafana_workspace_id=g-xxx`. a short lived API key to run the `apply` or `destroy` command.
Ensure you have necessary IAM permissions (`CreateWorkspaceApiKey, DeleteWorkspaceApiKey`)
7. Grafana API Key
- Give admin access to the SSO user you set up when creating the Amazon Managed Grafana Workspace:
- In the AWS Console, navigate to Amazon Grafana. In the left navigation bar, click **All workspaces**, then click on the workspace name you are using for this example.
- Under **Authentication** within **AWS Single Sign-On (SSO)**, click **Configure users and user groups**
- Check the box next to the SSO user you created and click **Make admin**
- From the workspace in the AWS console, click on the `Grafana workspace URL` to open the workspace
- If you don't see the gear icon in the left navigation bar, log out and log back in.
- Click on the gear icon, then click on the **API keys** tab.
- Click **Add API key**, fill in the _Key name_ field and select _Admin_ as the Role.
- Copy your API key into `terraform.tfvars` under the `grafana_api_key` variable (`grafana_api_key="xxx"`) or set as an environment variable on your CLI (`export TF_VAR_grafana_api_key="xxx"`)
```sh
export TF_VAR_grafana_api_key=`aws grafana create-workspace-api-key --key-name "observability-accelerator-$(date +%s)" --key-role ADMIN --seconds-to-live 1200 --workspace-id $TF_VAR_managed_grafana_workspace_id --query key --output text`
```
## Deploy ## Deploy
@@ -88,11 +77,31 @@ or if you had setup environment variables, run
terraform apply terraform apply
``` ```
## Additional configuration
For the purpose of the example, we have provided default values for some of the variables.
1. AWS Region
Specify the AWS Region where the resources will be deployed. Edit the `terraform.tfvars` file and modify `aws_region="..."`. You can also use environement variables `export TF_VAR_aws_region=xxx`.
2. Amazon Managed Service for Prometheus workspace
If you have an existing workspace, add `managed_prometheus_workspace_id=ws-xxx`
or use an environment variable `export TF_VAR_managed_prometheus_workspace_id=ws-xxx`.
## Visualization ## Visualization
1. Prometheus datasource on Grafana 1. Prometheus datasource on Grafana
Open your Grafana workspace and under Configuration -> Data sources, you should see `aws-observability-accelerator`. Open and click `Save & test`. You should see a notification confirming that the Amazon Managed Service for Prometheus workspace is ready to be used on Grafana. Make sure to open the link in the output. After a successful deployment, this will open
the Prometheus datasource configuration on Grafana.
Click `Save & test` and you should see a notification confirming that the Amazon Managed Service for Prometheus workspace is ready to be used on Grafana.
```bash
terraform output grafana_prometheus_datasource_test
```
2. Grafana dashboards 2. Grafana dashboards
@@ -113,7 +122,7 @@ Open the Amazon Managed Service for Prometheus console and view the details of y
To setup your alert receiver, with Amazon SNS, follow [this documentation](https://docs.aws.amazon.com/prometheus/latest/userguide/AMP-alertmanager-receiver.html) To setup your alert receiver, with Amazon SNS, follow [this documentation](https://docs.aws.amazon.com/prometheus/latest/userguide/AMP-alertmanager-receiver.html)
## Deploy an Example Application to Visualize ## Deploy an example application to visualize metrics
In this section we will deploy sample application and extract metrics using AWS OpenTelemetry collector In this section we will deploy sample application and extract metrics using AWS OpenTelemetry collector
@@ -161,7 +170,7 @@ kubectl apply -f -
kubectl get pods -n nginx-ingress-sample kubectl get pods -n nginx-ingress-sample
``` ```
#### Visualize the Application's dashboard #### Visualize the application's dashboard
Log back into your Managed Grafana Workspace and navigate to the dashboard side panel, click on `Observability Accelerator Dashboards` Folder and open the `NGINX` Dashboard. Log back into your Managed Grafana Workspace and navigate to the dashboard side panel, click on `Observability Accelerator Dashboards` Folder and open the `NGINX` Dashboard.
@@ -208,8 +217,8 @@ add this `managed_prometheus_region=xxx` and `managed_prometheus_workspace_id=ws
| Name | Source | Version | | Name | Source | Version |
|------|--------|---------| |------|--------|---------|
| <a name="module_eks_observability_accelerator"></a> [eks\_observability\_accelerator](#module\_eks\_observability\_accelerator) | ../../ | n/a | | <a name="module_aws_observability_accelerator"></a> [aws\_observability\_accelerator](#module\_aws\_observability\_accelerator) | ../../ | n/a |
| <a name="module_workloads_nginx"></a> [workloads\_nginx](#module\_workloads\_nginx) | ../../modules/workloads/nginx | n/a | | <a name="module_eks_monitoring"></a> [eks\_monitoring](#module\_eks\_monitoring) | ../../modules/eks-monitoring | n/a |
## Resources ## Resources
@@ -224,8 +233,8 @@ add this `managed_prometheus_region=xxx` and `managed_prometheus_workspace_id=ws
|------|-------------|------|---------|:--------:| |------|-------------|------|---------|:--------:|
| <a name="input_aws_region"></a> [aws\_region](#input\_aws\_region) | AWS Region | `string` | n/a | yes | | <a name="input_aws_region"></a> [aws\_region](#input\_aws\_region) | AWS Region | `string` | n/a | yes |
| <a name="input_eks_cluster_id"></a> [eks\_cluster\_id](#input\_eks\_cluster\_id) | EKS Cluster Id | `string` | n/a | yes | | <a name="input_eks_cluster_id"></a> [eks\_cluster\_id](#input\_eks\_cluster\_id) | EKS Cluster Id | `string` | n/a | yes |
| <a name="input_grafana_api_key"></a> [grafana\_api\_key](#input\_grafana\_api\_key) | API key for authorizing the Grafana provider to make changes to Amazon Managed Grafana | `string` | `""` | no | | <a name="input_grafana_api_key"></a> [grafana\_api\_key](#input\_grafana\_api\_key) | API key for authorizing the Grafana provider to make changes to Amazon Managed Grafana | `string` | n/a | yes |
| <a name="input_managed_grafana_workspace_id"></a> [managed\_grafana\_workspace\_id](#input\_managed\_grafana\_workspace\_id) | Amazon Managed Grafana (AMG) workspace ID | `string` | `""` | no | | <a name="input_managed_grafana_workspace_id"></a> [managed\_grafana\_workspace\_id](#input\_managed\_grafana\_workspace\_id) | Amazon Managed Grafana (AMG) workspace ID | `string` | n/a | yes |
| <a name="input_managed_prometheus_workspace_id"></a> [managed\_prometheus\_workspace\_id](#input\_managed\_prometheus\_workspace\_id) | Amazon Managed Service for Prometheus (AMP) workspace ID | `string` | `""` | no | | <a name="input_managed_prometheus_workspace_id"></a> [managed\_prometheus\_workspace\_id](#input\_managed\_prometheus\_workspace\_id) | Amazon Managed Service for Prometheus (AMP) workspace ID | `string` | `""` | no |
## Outputs ## Outputs
@@ -236,6 +245,7 @@ add this `managed_prometheus_region=xxx` and `managed_prometheus_workspace_id=ws
| <a name="output_eks_cluster_id"></a> [eks\_cluster\_id](#output\_eks\_cluster\_id) | EKS Cluster Id | | <a name="output_eks_cluster_id"></a> [eks\_cluster\_id](#output\_eks\_cluster\_id) | EKS Cluster Id |
| <a name="output_eks_cluster_version"></a> [eks\_cluster\_version](#output\_eks\_cluster\_version) | EKS Cluster version | | <a name="output_eks_cluster_version"></a> [eks\_cluster\_version](#output\_eks\_cluster\_version) | EKS Cluster version |
| <a name="output_grafana_dashboard_urls"></a> [grafana\_dashboard\_urls](#output\_grafana\_dashboard\_urls) | URLs for dashboards created | | <a name="output_grafana_dashboard_urls"></a> [grafana\_dashboard\_urls](#output\_grafana\_dashboard\_urls) | URLs for dashboards created |
| <a name="output_grafana_prometheus_datasource_test"></a> [grafana\_prometheus\_datasource\_test](#output\_grafana\_prometheus\_datasource\_test) | Grafana save & test URL for Amazon Managed Prometheus workspace |
| <a name="output_managed_prometheus_workspace_endpoint"></a> [managed\_prometheus\_workspace\_endpoint](#output\_managed\_prometheus\_workspace\_endpoint) | Amazon Managed Prometheus workspace endpoint | | <a name="output_managed_prometheus_workspace_endpoint"></a> [managed\_prometheus\_workspace\_endpoint](#output\_managed\_prometheus\_workspace\_endpoint) | Amazon Managed Prometheus workspace endpoint |
| <a name="output_managed_prometheus_workspace_id"></a> [managed\_prometheus\_workspace\_id](#output\_managed\_prometheus\_workspace\_id) | Amazon Managed Prometheus workspace ID | | <a name="output_managed_prometheus_workspace_id"></a> [managed\_prometheus\_workspace\_id](#output\_managed\_prometheus\_workspace\_id) | Amazon Managed Prometheus workspace ID |
<!-- END OF PRE-COMMIT-TERRAFORM DOCS HOOK --> <!-- END OF PRE-COMMIT-TERRAFORM DOCS HOOK -->
+18 -26
View File
@@ -25,8 +25,7 @@ provider "helm" {
} }
locals { locals {
region = var.aws_region region = var.aws_region
eks_cluster_endpoint = data.aws_eks_cluster.this.endpoint eks_cluster_endpoint = data.aws_eks_cluster.this.endpoint
create_new_workspace = var.managed_prometheus_workspace_id == "" ? true : false create_new_workspace = var.managed_prometheus_workspace_id == "" ? true : false
@@ -35,28 +34,17 @@ locals {
} }
} }
module "eks_observability_accelerator" { module "aws_observability_accelerator" {
# source = "aws-observability/terrarom-aws-observability-accelerator"
source = "../../" source = "../../"
# source = "github.com/aws-observability/terraform-aws-observability-accelerator?ref=v2.0.0"
aws_region = var.aws_region aws_region = var.aws_region
eks_cluster_id = var.eks_cluster_id
# deploys AWS Distro for OpenTelemetry operator into the cluster
enable_amazon_eks_adot = true
# reusing existing certificate manager? defaults to true
enable_cert_manager = true
# creates a new AMP workspace, defaults to true # creates a new AMP workspace, defaults to true
enable_managed_prometheus = local.create_new_workspace enable_managed_prometheus = local.create_new_workspace
# reusing existing AMP if specified # reusing existing AMP if specified
managed_prometheus_workspace_id = var.managed_prometheus_workspace_id managed_prometheus_workspace_id = var.managed_prometheus_workspace_id
managed_prometheus_workspace_region = null # defaults to the current region, useful for cross region scenarios (same account)
# sets up the AMP alert manager at the workspace level
enable_alertmanager = true
# reusing existing Amazon Managed Grafana workspace # reusing existing Amazon Managed Grafana workspace
enable_managed_grafana = false enable_managed_grafana = false
@@ -67,24 +55,28 @@ module "eks_observability_accelerator" {
} }
provider "grafana" { provider "grafana" {
url = module.eks_observability_accelerator.managed_grafana_workspace_endpoint url = module.aws_observability_accelerator.managed_grafana_workspace_endpoint
auth = var.grafana_api_key auth = var.grafana_api_key
} }
module "workloads_nginx" { module "eks_monitoring" {
source = "../../modules/workloads/nginx" source = "../../modules/eks-monitoring"
# source = "github.com/aws-observability/terraform-aws-observability-accelerator//modules/eks-monitoring?ref=v2.0.0"
eks_cluster_id = module.eks_observability_accelerator.eks_cluster_id # enable NGINX metrics collection, dashboards and alerts rules creation
enable_nginx = true
dashboards_folder_id = module.eks_observability_accelerator.grafana_dashboards_folder_id eks_cluster_id = var.eks_cluster_id
managed_prometheus_workspace_id = module.eks_observability_accelerator.managed_prometheus_workspace_id
managed_prometheus_workspace_endpoint = module.eks_observability_accelerator.managed_prometheus_workspace_endpoint dashboards_folder_id = module.aws_observability_accelerator.grafana_dashboards_folder_id
managed_prometheus_workspace_region = module.eks_observability_accelerator.managed_prometheus_workspace_region managed_prometheus_workspace_id = module.aws_observability_accelerator.managed_prometheus_workspace_id
managed_prometheus_workspace_endpoint = module.aws_observability_accelerator.managed_prometheus_workspace_endpoint
managed_prometheus_workspace_region = module.aws_observability_accelerator.managed_prometheus_workspace_region
tags = local.tags tags = local.tags
depends_on = [ depends_on = [
module.eks_observability_accelerator module.aws_observability_accelerator
] ]
} }
+19 -14
View File
@@ -1,29 +1,34 @@
output "eks_cluster_id" {
description = "EKS Cluster Id"
value = module.eks_observability_accelerator.eks_cluster_id
}
output "aws_region" { output "aws_region" {
description = "AWS Region" description = "AWS Region"
value = module.eks_observability_accelerator.aws_region value = module.aws_observability_accelerator.aws_region
}
output "eks_cluster_version" {
description = "EKS Cluster version"
value = module.eks_observability_accelerator.eks_cluster_version
} }
output "managed_prometheus_workspace_endpoint" { output "managed_prometheus_workspace_endpoint" {
description = "Amazon Managed Prometheus workspace endpoint" description = "Amazon Managed Prometheus workspace endpoint"
value = module.eks_observability_accelerator.managed_prometheus_workspace_endpoint value = module.aws_observability_accelerator.managed_prometheus_workspace_endpoint
} }
output "managed_prometheus_workspace_id" { output "managed_prometheus_workspace_id" {
description = "Amazon Managed Prometheus workspace ID" description = "Amazon Managed Prometheus workspace ID"
value = module.eks_observability_accelerator.managed_prometheus_workspace_id value = module.aws_observability_accelerator.managed_prometheus_workspace_id
} }
output "grafana_dashboard_urls" { output "grafana_dashboard_urls" {
description = "URLs for dashboards created" description = "URLs for dashboards created"
value = module.workloads_nginx.grafana_dashboard_urls value = module.eks_monitoring.grafana_dashboard_urls
}
output "grafana_prometheus_datasource_test" {
description = "Grafana save & test URL for Amazon Managed Prometheus workspace"
value = module.aws_observability_accelerator.grafana_prometheus_datasource_test
}
output "eks_cluster_version" {
description = "EKS Cluster version"
value = module.eks_monitoring.eks_cluster_version
}
output "eks_cluster_id" {
description = "EKS Cluster Id"
value = module.eks_monitoring.eks_cluster_id
} }
@@ -14,11 +14,9 @@ variable "managed_prometheus_workspace_id" {
variable "managed_grafana_workspace_id" { variable "managed_grafana_workspace_id" {
description = "Amazon Managed Grafana (AMG) workspace ID" description = "Amazon Managed Grafana (AMG) workspace ID"
type = string type = string
default = ""
} }
variable "grafana_api_key" { variable "grafana_api_key" {
description = "API key for authorizing the Grafana provider to make changes to Amazon Managed Grafana" description = "API key for authorizing the Grafana provider to make changes to Amazon Managed Grafana"
type = string type = string
default = ""
sensitive = true sensitive = true
} }
@@ -1,17 +1,16 @@
# Existing Cluster with the AWS Observability accelerator base module and Infrastructure monitoring # Existing Cluster with the AWS Observability accelerator base module and Infrastructure monitoring
This example demonstrates how to use the AWS Observability Accelerator Terraform This example demonstrates how to use the AWS Observability Accelerator Terraform
modules with Infrastructure monitoring enabled. modules with Infrastructure monitoring enabled.
The current example deploys the [AWS Distro for OpenTelemetry Operator](https://docs.aws.amazon.com/eks/latest/userguide/opentelemetry.html) for Amazon EKS with its requirements and make use of existing The current example deploys the [AWS Distro for OpenTelemetry Operator](https://docs.aws.amazon.com/eks/latest/userguide/opentelemetry.html)
Amazon Managed Service for Prometheus and Amazon Managed Grafana workspaces. for Amazon EKS with its requirements and make use of an existing Amazon Managed Grafana workspace.
It creates a new Amazon Managed Service for Prometheus workspace unless provided with an existing one to reuse.
It is based on the `infrastructure monitoring`, one of our [workloads modules](../../modules/workloads/) It uses the `EKS monitoring` [module](../../modules/eks-monitoring/)
to provide an existing EKS cluster with an OpenTelemetry collector, to provide an existing EKS cluster with an OpenTelemetry collector,
curated Grafana dashboards, Prometheus alerting and recording rules with multiple curated Grafana dashboards, Prometheus alerting and recording rules with multiple
configuration options on the cluster infrastructure. configuration options on the cluster infrastructure.
## Prerequisites ## Prerequisites
Ensure that you have the following tools installed locally: Ensure that you have the following tools installed locally:
@@ -20,7 +19,6 @@ Ensure that you have the following tools installed locally:
2. [kubectl](https://kubernetes.io/docs/tasks/tools/) 2. [kubectl](https://kubernetes.io/docs/tasks/tools/)
3. [terraform](https://learn.hashicorp.com/tutorials/terraform/install-cli) 3. [terraform](https://learn.hashicorp.com/tutorials/terraform/install-cli)
## Setup ## Setup
This example uses a local terraform state. If you need states to be saved remotely, This example uses a local terraform state. If you need states to be saved remotely,
@@ -39,11 +37,7 @@ cd examples/existing-cluster-with-base-and-infra
terraform init terraform init
``` ```
3. AWS Region 3. Amazon EKS Cluster
Specify the AWS Region where the resources will be deployed. Edit the `terraform.tfvars` file and modify `aws_region="..."`. You can also use environement variables `export TF_VAR_aws_region=xxx`.
4. Amazon EKS Cluster
To run this example, you need to provide your EKS cluster name. To run this example, you need to provide your EKS cluster name.
If you don't have a cluster ready, visit [this example](../eks-cluster-with-vpc) If you don't have a cluster ready, visit [this example](../eks-cluster-with-vpc)
@@ -51,24 +45,15 @@ first to create a new one.
Add your cluster name for `eks_cluster_id="..."` to the `terraform.tfvars` or use an environment variable `export TF_VAR_eks_cluster_id=xxx`. Add your cluster name for `eks_cluster_id="..."` to the `terraform.tfvars` or use an environment variable `export TF_VAR_eks_cluster_id=xxx`.
5. Amazon Managed Service for Prometheus workspace (optional) 4. Amazon Managed Grafana workspace
If you have an existing workspace, add `managed_prometheus_workspace_id=ws-xxx`
or use an environment variable `export TF_VAR_managed_prometheus_workspace_id=ws-xxx`.
If you don't specify anything a new workspace will be created for you.
6. Amazon Managed Grafana workspace
To run this example you need an Amazon Managed Grafana workspace. If you have an existing workspace, create an environment variable `export TF_VAR_managed_grafana_workspace_id=g-xxx`. To run this example you need an Amazon Managed Grafana workspace. If you have an existing workspace, create an environment variable `export TF_VAR_managed_grafana_workspace_id=g-xxx`.
To create a new one, visit our Amazon Managed Grafana [documentation](https://docs.aws.amazon.com/grafana/latest/userguide/getting-started-with-AMG.html).
Make sure to provide the workspace with Amazon Managed Service for Prometheus read permissions.
To create a new one, within this example's Terraform state (sharing the same lifecycle with all the other resources): > In the URL `https://g-xyz.grafana-workspace.eu-central-1.amazonaws.com`, the workspace ID would be `g-xyz`
- Edit main.tf and set `enable_managed_grafana = true` 5. <a name="apikey"></a> Grafana API Key
- Run `terraform apply -target "module.eks_observability_accelerator.module.managed_grafana[0].aws_grafana_workspace.this[0]"`.
- Run `export TF_VAR_managed_grafana_workspace_id=$(terraform output --raw managed_grafana_workspace_id)`.
7. <a name="apikey"></a> Grafana API Key
Amazon Managed Service for Grafana provides a control plane API for generating Grafana API keys. We will provide to Terraform Amazon Managed Service for Grafana provides a control plane API for generating Grafana API keys. We will provide to Terraform
a short lived API key to run the `apply` or `destroy` command. a short lived API key to run the `apply` or `destroy` command.
@@ -90,11 +75,31 @@ or if you had only setup environment variables, run
terraform apply terraform apply
``` ```
## Additional configuration
For the purpose of the example, we have provided default values for some of the variables.
1. AWS Region
Specify the AWS Region where the resources will be deployed. Edit the `terraform.tfvars` file and modify `aws_region="..."`. You can also use environement variables `export TF_VAR_aws_region=xxx`.
2. Amazon Managed Service for Prometheus workspace
If you have an existing workspace, add `managed_prometheus_workspace_id=ws-xxx`
or use an environment variable `export TF_VAR_managed_prometheus_workspace_id=ws-xxx`.
## Visualization ## Visualization
1. Prometheus datasource on Grafana 1. Prometheus datasource on Grafana
Open your Grafana workspace and under Configuration -> Data sources, you should see `aws-observability-accelerator`. Open and click `Save & test`. You should see a notification confirming that the Amazon Managed Service for Prometheus workspace is ready to be used on Grafana. Make sure to open the link in the output. After a successful deployment, this will open
the Prometheus datasource configuration on Grafana.
Click `Save & test` and you should see a notification confirming that the Amazon Managed Service for Prometheus workspace is ready to be used on Grafana.
```bash
terraform output grafana_prometheus_datasource_test
```
2. Grafana dashboards 2. Grafana dashboards
@@ -115,17 +120,6 @@ Open the Amazon Managed Service for Prometheus console and view the details of y
To setup your alert receiver, with Amazon SNS, follow [this documentation](https://docs.aws.amazon.com/prometheus/latest/userguide/AMP-alertmanager-receiver.html) To setup your alert receiver, with Amazon SNS, follow [this documentation](https://docs.aws.amazon.com/prometheus/latest/userguide/AMP-alertmanager-receiver.html)
## Advanced configuration
1. Cross-region Amazon Managed Prometheus workspace
If your existing Amazon Managed Prometheus workspace is in another AWS Region,
add this `managed_prometheus_region=xxx` and `managed_prometheus_workspace_id=ws-xxx`.
2. Cross-region Amazon Managed Grafana workspace
If your existing Amazon Managed Prometheus workspace is in another AWS Region,
add this `managed_prometheus_region=xxx` and `managed_prometheus_workspace_id=ws-xxx`.
## Destroy resources ## Destroy resources
@@ -159,8 +153,8 @@ terraform destroy -var-file=terraform.tfvars
| Name | Source | Version | | Name | Source | Version |
|------|--------|---------| |------|--------|---------|
| <a name="module_eks_observability_accelerator"></a> [eks\_observability\_accelerator](#module\_eks\_observability\_accelerator) | ../../ | n/a | | <a name="module_aws_observability_accelerator"></a> [aws\_observability\_accelerator](#module\_aws\_observability\_accelerator) | ../../ | n/a |
| <a name="module_workloads_infra"></a> [workloads\_infra](#module\_workloads\_infra) | ../../modules/workloads/infra | n/a | | <a name="module_eks_monitoring"></a> [eks\_monitoring](#module\_eks\_monitoring) | ../../modules/eks-monitoring | n/a |
## Resources ## Resources
@@ -175,8 +169,8 @@ terraform destroy -var-file=terraform.tfvars
|------|-------------|------|---------|:--------:| |------|-------------|------|---------|:--------:|
| <a name="input_aws_region"></a> [aws\_region](#input\_aws\_region) | AWS Region | `string` | n/a | yes | | <a name="input_aws_region"></a> [aws\_region](#input\_aws\_region) | AWS Region | `string` | n/a | yes |
| <a name="input_eks_cluster_id"></a> [eks\_cluster\_id](#input\_eks\_cluster\_id) | Name of the EKS cluster | `string` | n/a | yes | | <a name="input_eks_cluster_id"></a> [eks\_cluster\_id](#input\_eks\_cluster\_id) | Name of the EKS cluster | `string` | n/a | yes |
| <a name="input_grafana_api_key"></a> [grafana\_api\_key](#input\_grafana\_api\_key) | API key for authorizing the Grafana provider to make changes to Amazon Managed Grafana | `string` | `""` | no | | <a name="input_grafana_api_key"></a> [grafana\_api\_key](#input\_grafana\_api\_key) | API key for authorizing the Grafana provider to make changes to Amazon Managed Grafana | `string` | n/a | yes |
| <a name="input_managed_grafana_workspace_id"></a> [managed\_grafana\_workspace\_id](#input\_managed\_grafana\_workspace\_id) | Amazon Managed Grafana Workspace ID | `string` | `""` | no | | <a name="input_managed_grafana_workspace_id"></a> [managed\_grafana\_workspace\_id](#input\_managed\_grafana\_workspace\_id) | Amazon Managed Grafana Workspace ID | `string` | n/a | yes |
| <a name="input_managed_prometheus_workspace_id"></a> [managed\_prometheus\_workspace\_id](#input\_managed\_prometheus\_workspace\_id) | Amazon Managed Service for Prometheus Workspace ID | `string` | `""` | no | | <a name="input_managed_prometheus_workspace_id"></a> [managed\_prometheus\_workspace\_id](#input\_managed\_prometheus\_workspace\_id) | Amazon Managed Service for Prometheus Workspace ID | `string` | `""` | no |
## Outputs ## Outputs
@@ -187,6 +181,7 @@ terraform destroy -var-file=terraform.tfvars
| <a name="output_eks_cluster_id"></a> [eks\_cluster\_id](#output\_eks\_cluster\_id) | EKS Cluster Id | | <a name="output_eks_cluster_id"></a> [eks\_cluster\_id](#output\_eks\_cluster\_id) | EKS Cluster Id |
| <a name="output_eks_cluster_version"></a> [eks\_cluster\_version](#output\_eks\_cluster\_version) | EKS Cluster version | | <a name="output_eks_cluster_version"></a> [eks\_cluster\_version](#output\_eks\_cluster\_version) | EKS Cluster version |
| <a name="output_grafana_dashboard_urls"></a> [grafana\_dashboard\_urls](#output\_grafana\_dashboard\_urls) | URLs for dashboards created | | <a name="output_grafana_dashboard_urls"></a> [grafana\_dashboard\_urls](#output\_grafana\_dashboard\_urls) | URLs for dashboards created |
| <a name="output_grafana_prometheus_datasource_test"></a> [grafana\_prometheus\_datasource\_test](#output\_grafana\_prometheus\_datasource\_test) | Grafana save & test URL for Amazon Managed Prometheus workspace |
| <a name="output_managed_grafana_workspace_id"></a> [managed\_grafana\_workspace\_id](#output\_managed\_grafana\_workspace\_id) | Amazon Managed Grafana workspace ID | | <a name="output_managed_grafana_workspace_id"></a> [managed\_grafana\_workspace\_id](#output\_managed\_grafana\_workspace\_id) | Amazon Managed Grafana workspace ID |
| <a name="output_managed_prometheus_workspace_endpoint"></a> [managed\_prometheus\_workspace\_endpoint](#output\_managed\_prometheus\_workspace\_endpoint) | Amazon Managed Prometheus workspace endpoint | | <a name="output_managed_prometheus_workspace_endpoint"></a> [managed\_prometheus\_workspace\_endpoint](#output\_managed\_prometheus\_workspace\_endpoint) | Amazon Managed Prometheus workspace endpoint |
| <a name="output_managed_prometheus_workspace_id"></a> [managed\_prometheus\_workspace\_id](#output\_managed\_prometheus\_workspace\_id) | Amazon Managed Prometheus workspace ID | | <a name="output_managed_prometheus_workspace_id"></a> [managed\_prometheus\_workspace\_id](#output\_managed\_prometheus\_workspace\_id) | Amazon Managed Prometheus workspace ID |
@@ -34,25 +34,17 @@ locals {
} }
# deploys the base module # deploys the base module
module "eks_observability_accelerator" { module "aws_observability_accelerator" {
# source = "aws-observability/terrarom-aws-observability-accelerator"
source = "../../" source = "../../"
# source = "github.com/aws-observability/terraform-aws-observability-accelerator?ref=v2.0.0"
aws_region = var.aws_region aws_region = var.aws_region
eks_cluster_id = var.eks_cluster_id
# deploys AWS Distro for OpenTelemetry operator into the cluster
enable_amazon_eks_adot = true
# reusing existing certificate manager? defaults to true
enable_cert_manager = true
# creates a new Amazon Managed Prometheus workspace, defaults to true # creates a new Amazon Managed Prometheus workspace, defaults to true
enable_managed_prometheus = local.create_new_workspace enable_managed_prometheus = local.create_new_workspace
# reusing existing Amazon Managed Prometheus if specified # reusing existing Amazon Managed Prometheus if specified
managed_prometheus_workspace_id = var.managed_prometheus_workspace_id managed_prometheus_workspace_id = var.managed_prometheus_workspace_id
managed_prometheus_workspace_region = null # defaults to the current region, useful for cross region scenarios (same account)
# sets up the Amazon Managed Prometheus alert manager at the workspace level # sets up the Amazon Managed Prometheus alert manager at the workspace level
enable_alertmanager = true enable_alertmanager = true
@@ -70,20 +62,27 @@ module "eks_observability_accelerator" {
# any provider blocks. # any provider blocks.
# This allows forcing dependency between base and workloads module # This allows forcing dependency between base and workloads module
provider "grafana" { provider "grafana" {
url = module.eks_observability_accelerator.managed_grafana_workspace_endpoint url = module.aws_observability_accelerator.managed_grafana_workspace_endpoint
auth = var.grafana_api_key auth = var.grafana_api_key
} }
module "workloads_infra" { module "eks_monitoring" {
source = "../../modules/workloads/infra" source = "../../modules/eks-monitoring"
# source = "github.com/aws-observability/terraform-aws-observability-accelerator//modules/eks-monitoring?ref=v2.0.0"
eks_cluster_id = module.eks_observability_accelerator.eks_cluster_id eks_cluster_id = var.eks_cluster_id
dashboards_folder_id = module.eks_observability_accelerator.grafana_dashboards_folder_id # deploys AWS Distro for OpenTelemetry operator into the cluster
managed_prometheus_workspace_id = module.eks_observability_accelerator.managed_prometheus_workspace_id enable_amazon_eks_adot = true
managed_prometheus_workspace_endpoint = module.eks_observability_accelerator.managed_prometheus_workspace_endpoint # reusing existing certificate manager? defaults to true
managed_prometheus_workspace_region = module.eks_observability_accelerator.managed_prometheus_workspace_region enable_cert_manager = true
dashboards_folder_id = module.aws_observability_accelerator.grafana_dashboards_folder_id
managed_prometheus_workspace_id = module.aws_observability_accelerator.managed_prometheus_workspace_id
managed_prometheus_workspace_endpoint = module.aws_observability_accelerator.managed_prometheus_workspace_endpoint
managed_prometheus_workspace_region = module.aws_observability_accelerator.managed_prometheus_workspace_region
# optional, defaults to 60s interval and 15s timeout # optional, defaults to 60s interval and 15s timeout
prometheus_config = { prometheus_config = {
@@ -94,6 +93,6 @@ module "workloads_infra" {
tags = local.tags tags = local.tags
depends_on = [ depends_on = [
module.eks_observability_accelerator module.aws_observability_accelerator
] ]
} }
@@ -1,34 +1,38 @@
output "eks_cluster_id" {
description = "EKS Cluster Id"
value = module.eks_observability_accelerator.eks_cluster_id
}
output "aws_region" { output "aws_region" {
description = "AWS Region" description = "AWS Region"
value = module.eks_observability_accelerator.aws_region value = module.aws_observability_accelerator.aws_region
}
output "eks_cluster_version" {
description = "EKS Cluster version"
value = module.eks_observability_accelerator.eks_cluster_version
} }
output "managed_prometheus_workspace_endpoint" { output "managed_prometheus_workspace_endpoint" {
description = "Amazon Managed Prometheus workspace endpoint" description = "Amazon Managed Prometheus workspace endpoint"
value = module.eks_observability_accelerator.managed_prometheus_workspace_endpoint value = module.aws_observability_accelerator.managed_prometheus_workspace_endpoint
} }
output "managed_prometheus_workspace_id" { output "managed_prometheus_workspace_id" {
description = "Amazon Managed Prometheus workspace ID" description = "Amazon Managed Prometheus workspace ID"
value = module.eks_observability_accelerator.managed_prometheus_workspace_id value = module.aws_observability_accelerator.managed_prometheus_workspace_id
} }
output "managed_grafana_workspace_id" { output "managed_grafana_workspace_id" {
description = "Amazon Managed Grafana workspace ID" description = "Amazon Managed Grafana workspace ID"
value = module.eks_observability_accelerator.managed_grafana_workspace_id value = module.aws_observability_accelerator.managed_grafana_workspace_id
} }
output "grafana_dashboard_urls" { output "grafana_dashboard_urls" {
description = "URLs for dashboards created" description = "URLs for dashboards created"
value = module.workloads_infra.grafana_dashboard_urls value = module.eks_monitoring.grafana_dashboard_urls
}
output "grafana_prometheus_datasource_test" {
description = "Grafana save & test URL for Amazon Managed Prometheus workspace"
value = module.aws_observability_accelerator.grafana_prometheus_datasource_test
}
output "eks_cluster_version" {
description = "EKS Cluster version"
value = module.eks_monitoring.eks_cluster_version
}
output "eks_cluster_id" {
description = "EKS Cluster Id"
value = module.eks_monitoring.eks_cluster_id
} }
@@ -1,8 +1,8 @@
# (mandatory) AWS Region where your resources will be located # (mandatory) AWS Region where your resources will be located
aws_region = "" aws_region = "eu-central-1"
# (mandatory) EKS Cluster name # (mandatory) EKS Cluster name
eks_cluster_id = "" eks_cluster_id = "eks-cluster-with-vpc"
# (optional) Leave it empty for a new workspace to be created # (optional) Leave it empty for a new workspace to be created
managed_prometheus_workspace_id = "" managed_prometheus_workspace_id = ""
@@ -14,11 +14,9 @@ variable "managed_prometheus_workspace_id" {
variable "managed_grafana_workspace_id" { variable "managed_grafana_workspace_id" {
description = "Amazon Managed Grafana Workspace ID" description = "Amazon Managed Grafana Workspace ID"
type = string type = string
default = ""
} }
variable "grafana_api_key" { variable "grafana_api_key" {
description = "API key for authorizing the Grafana provider to make changes to Amazon Managed Grafana" description = "API key for authorizing the Grafana provider to make changes to Amazon Managed Grafana"
type = string type = string
default = ""
sensitive = true sensitive = true
} }
@@ -22,7 +22,7 @@ resource "grafana_folder" "this" {
} }
module "managed_prometheus_monitoring" { module "managed_prometheus_monitoring" {
source = "../../modules/workloads/managed-prometheus-monitoring" source = "../../modules/managed-prometheus-monitoring"
dashboards_folder_id = resource.grafana_folder.this.id dashboards_folder_id = resource.grafana_folder.this.id
aws_region = local.region aws_region = local.region
managed_prometheus_workspace_ids = var.managed_prometheus_workspace_ids managed_prometheus_workspace_ids = var.managed_prometheus_workspace_ids
-26
View File
@@ -1,23 +1,11 @@
data "aws_partition" "current" {}
data "aws_caller_identity" "current" {}
data "aws_region" "current" {} data "aws_region" "current" {}
data "aws_eks_cluster" "eks_cluster" {
name = var.eks_cluster_id
}
data "aws_grafana_workspace" "this" { data "aws_grafana_workspace" "this" {
count = var.managed_grafana_workspace_id == "" ? 0 : 1 count = var.managed_grafana_workspace_id == "" ? 0 : 1
workspace_id = var.managed_grafana_workspace_id workspace_id = var.managed_grafana_workspace_id
} }
locals { locals {
eks_oidc_issuer_url = replace(data.aws_eks_cluster.eks_cluster.identity[0].oidc[0].issuer, "https://", "")
eks_cluster_endpoint = data.aws_eks_cluster.eks_cluster.endpoint
eks_cluster_version = data.aws_eks_cluster.eks_cluster.version
# if region is not passed, we assume the current one # if region is not passed, we assume the current one
amp_ws_region = coalesce(var.managed_prometheus_workspace_region, data.aws_region.current.name) amp_ws_region = coalesce(var.managed_prometheus_workspace_region, data.aws_region.current.name)
amp_ws_id = var.enable_managed_prometheus ? aws_prometheus_workspace.this[0].id : var.managed_prometheus_workspace_id amp_ws_id = var.enable_managed_prometheus ? aws_prometheus_workspace.this[0].id : var.managed_prometheus_workspace_id
@@ -28,19 +16,5 @@ locals {
amg_ws_endpoint = var.managed_grafana_workspace_id == "" ? "https://${module.managed_grafana[0].workspace_endpoint}" : "https://${data.aws_grafana_workspace.this[0].endpoint}" amg_ws_endpoint = var.managed_grafana_workspace_id == "" ? "https://${module.managed_grafana[0].workspace_endpoint}" : "https://${data.aws_grafana_workspace.this[0].endpoint}"
amg_ws_id = var.managed_grafana_workspace_id == "" ? split(".", module.managed_grafana[0].workspace_endpoint)[0] : var.managed_grafana_workspace_id amg_ws_id = var.managed_grafana_workspace_id == "" ? split(".", module.managed_grafana[0].workspace_endpoint)[0] : var.managed_grafana_workspace_id
context = {
aws_caller_identity_account_id = data.aws_caller_identity.current.account_id
aws_caller_identity_arn = data.aws_caller_identity.current.arn
aws_eks_cluster_endpoint = local.eks_cluster_endpoint
aws_partition_id = data.aws_partition.current.partition
aws_region_name = data.aws_region.current.name
eks_cluster_id = var.eks_cluster_id
eks_oidc_issuer_url = local.eks_oidc_issuer_url
eks_oidc_provider_arn = "arn:${data.aws_partition.current.partition}:iam::${data.aws_caller_identity.current.account_id}:oidc-provider/${local.eks_oidc_issuer_url}"
tags = var.tags
irsa_iam_role_path = var.irsa_iam_role_path
irsa_iam_permissions_boundary = var.irsa_iam_permissions_boundary
}
name = "aws-observability-accelerator" name = "aws-observability-accelerator"
} }
-9
View File
@@ -1,12 +1,3 @@
module "operator" {
source = "./modules/add-ons/adot-operator"
count = var.enable_amazon_eks_adot ? 1 : 0
enable_cert_manager = var.enable_cert_manager
kubernetes_version = local.eks_cluster_version
addon_context = local.context
}
resource "aws_prometheus_workspace" "this" { resource "aws_prometheus_workspace" "this" {
count = var.enable_managed_prometheus ? 1 : 0 count = var.enable_managed_prometheus ? 1 : 0
+4 -4
View File
@@ -27,14 +27,14 @@ nav:
- Concepts: concepts.md - Concepts: concepts.md
- Amazon EKS: - Amazon EKS:
- Infrastructure monitoring: eks/index.md - Infrastructure monitoring: eks/index.md
- Java/JMX: eks/java.md
- Nginx: eks/nginx.md
- Teardown: eks/destroy.md - Teardown: eks/destroy.md
- Workload Monitoring: - Monitoring Managed Service for Prometheus Workspaces: workloads/managed-prometheus.md
- Java/JMX: workloads/java.md
- Nginx: workloads/nginx.md
- Amazon Managed Service for Prometheus Workspaces: workloads/managed-prometheus.md
- Supporting Examples: - Supporting Examples:
- EKS Cluster with VPC: helpers/new-eks-cluster.md - EKS Cluster with VPC: helpers/new-eks-cluster.md
# - Amazon Managed Grafana setup: helpers/managed-grafana.md # - Amazon Managed Grafana setup: helpers/managed-grafana.md
- Support & feedback: support.md
- Contributors: contributors.md - Contributors: contributors.md
markdown_extensions: markdown_extensions:
@@ -33,6 +33,9 @@ This module is inspired from the open source [kube-prometheus-stack](https://git
| Name | Source | Version | | Name | Source | Version |
|------|--------|---------| |------|--------|---------|
| <a name="module_helm_addon"></a> [helm\_addon](#module\_helm\_addon) | github.com/aws-ia/terraform-aws-eks-blueprints//modules/kubernetes-addons/helm-addon | v4.13.1 | | <a name="module_helm_addon"></a> [helm\_addon](#module\_helm\_addon) | github.com/aws-ia/terraform-aws-eks-blueprints//modules/kubernetes-addons/helm-addon | v4.13.1 |
| <a name="module_java_monitoring"></a> [java\_monitoring](#module\_java\_monitoring) | ./patterns/java | n/a |
| <a name="module_nginx_monitoring"></a> [nginx\_monitoring](#module\_nginx\_monitoring) | ./patterns/nginx | n/a |
| <a name="module_operator"></a> [operator](#module\_operator) | ./add-ons/adot-operator | n/a |
## Resources ## Resources
@@ -61,20 +64,25 @@ This module is inspired from the open source [kube-prometheus-stack](https://git
| <a name="input_dashboards_folder_id"></a> [dashboards\_folder\_id](#input\_dashboards\_folder\_id) | Grafana folder ID for automatic dashboards | `string` | n/a | yes | | <a name="input_dashboards_folder_id"></a> [dashboards\_folder\_id](#input\_dashboards\_folder\_id) | Grafana folder ID for automatic dashboards | `string` | n/a | yes |
| <a name="input_eks_cluster_id"></a> [eks\_cluster\_id](#input\_eks\_cluster\_id) | EKS Cluster Id | `string` | n/a | yes | | <a name="input_eks_cluster_id"></a> [eks\_cluster\_id](#input\_eks\_cluster\_id) | EKS Cluster Id | `string` | n/a | yes |
| <a name="input_enable_alerting_rules"></a> [enable\_alerting\_rules](#input\_enable\_alerting\_rules) | Enables or disables Managed Prometheus alerting rules | `bool` | `true` | no | | <a name="input_enable_alerting_rules"></a> [enable\_alerting\_rules](#input\_enable\_alerting\_rules) | Enables or disables Managed Prometheus alerting rules | `bool` | `true` | no |
| <a name="input_enable_amazon_eks_adot"></a> [enable\_amazon\_eks\_adot](#input\_enable\_amazon\_eks\_adot) | Enables the ADOT Operator on the EKS Cluster | `bool` | `true` | no |
| <a name="input_enable_cert_manager"></a> [enable\_cert\_manager](#input\_enable\_cert\_manager) | Allow reusing an existing installation of cert-manager | `bool` | `true` | no |
| <a name="input_enable_custom_metrics"></a> [enable\_custom\_metrics](#input\_enable\_custom\_metrics) | Allows additional metrics collection for config elements in the `custom_metrics_config` config object. Automatic dashboards are not included | `bool` | `false` | no | | <a name="input_enable_custom_metrics"></a> [enable\_custom\_metrics](#input\_enable\_custom\_metrics) | Allows additional metrics collection for config elements in the `custom_metrics_config` config object. Automatic dashboards are not included | `bool` | `false` | no |
| <a name="input_enable_dashboards"></a> [enable\_dashboards](#input\_enable\_dashboards) | Enables or disables curated dashboards | `bool` | `true` | no | | <a name="input_enable_dashboards"></a> [enable\_dashboards](#input\_enable\_dashboards) | Enables or disables curated dashboards | `bool` | `true` | no |
| <a name="input_enable_java"></a> [enable\_java](#input\_enable\_java) | Enable Java workloads monitoring, alerting and default dashboards | `bool` | `false` | no |
| <a name="input_enable_kube_state_metrics"></a> [enable\_kube\_state\_metrics](#input\_enable\_kube\_state\_metrics) | Enables or disables Kube State metrics exporter. Disabling this might affect some data in the dashboards | `bool` | `true` | no | | <a name="input_enable_kube_state_metrics"></a> [enable\_kube\_state\_metrics](#input\_enable\_kube\_state\_metrics) | Enables or disables Kube State metrics exporter. Disabling this might affect some data in the dashboards | `bool` | `true` | no |
| <a name="input_enable_nginx"></a> [enable\_nginx](#input\_enable\_nginx) | Enable NGINX workloads monitoring, alerting and default dashboards | `bool` | `false` | no |
| <a name="input_enable_node_exporter"></a> [enable\_node\_exporter](#input\_enable\_node\_exporter) | Enables or disables Node exporter. Disabling this might affect some data in the dashboards | `bool` | `true` | no | | <a name="input_enable_node_exporter"></a> [enable\_node\_exporter](#input\_enable\_node\_exporter) | Enables or disables Node exporter. Disabling this might affect some data in the dashboards | `bool` | `true` | no |
| <a name="input_enable_recording_rules"></a> [enable\_recording\_rules](#input\_enable\_recording\_rules) | Enables or disables Managed Prometheus recording rules. Disabling this might affect some data in the dashboards | `bool` | `true` | no |
| <a name="input_enable_tracing"></a> [enable\_tracing](#input\_enable\_tracing) | (Experimental) Enables tracing with AWS X-Ray. This changes the deploy mode of the collector to daemon set. Requirement: adot add-on <= 0.58-build.0 | `bool` | `false` | no | | <a name="input_enable_tracing"></a> [enable\_tracing](#input\_enable\_tracing) | (Experimental) Enables tracing with AWS X-Ray. This changes the deploy mode of the collector to daemon set. Requirement: adot add-on <= 0.58-build.0 | `bool` | `false` | no |
| <a name="input_helm_config"></a> [helm\_config](#input\_helm\_config) | Helm Config for Prometheus | `any` | `{}` | no | | <a name="input_helm_config"></a> [helm\_config](#input\_helm\_config) | Helm Config for Prometheus | `any` | `{}` | no |
| <a name="input_irsa_iam_permissions_boundary"></a> [irsa\_iam\_permissions\_boundary](#input\_irsa\_iam\_permissions\_boundary) | IAM permissions boundary for IRSA roles | `string` | `null` | no | | <a name="input_irsa_iam_permissions_boundary"></a> [irsa\_iam\_permissions\_boundary](#input\_irsa\_iam\_permissions\_boundary) | IAM permissions boundary for IRSA roles | `string` | `null` | no |
| <a name="input_irsa_iam_role_path"></a> [irsa\_iam\_role\_path](#input\_irsa\_iam\_role\_path) | IAM role path for IRSA roles | `string` | `"/"` | no | | <a name="input_irsa_iam_role_path"></a> [irsa\_iam\_role\_path](#input\_irsa\_iam\_role\_path) | IAM role path for IRSA roles | `string` | `"/"` | no |
| <a name="input_java_config"></a> [java\_config](#input\_java\_config) | Configuration object for Java/JMX monitoring | <pre>object({<br> enable_alerting_rules = bool<br> scrape_sample_limit = number<br> })</pre> | <pre>{<br> "enable_alerting_rules": true,<br> "scrape_sample_limit": 1000<br>}</pre> | no |
| <a name="input_ksm_config"></a> [ksm\_config](#input\_ksm\_config) | Kube State metrics configuration | <pre>object({<br> create_namespace = bool<br> k8s_namespace = string<br> helm_chart_name = string<br> helm_chart_version = string<br> helm_release_name = string<br> helm_repo_url = string<br> helm_settings = map(string)<br> helm_values = map(any)<br><br> scrape_interval = string<br> scrape_timeout = string<br> })</pre> | <pre>{<br> "create_namespace": true,<br> "helm_chart_name": "kube-state-metrics",<br> "helm_chart_version": "4.24.0",<br> "helm_release_name": "kube-state-metrics",<br> "helm_repo_url": "https://prometheus-community.github.io/helm-charts",<br> "helm_settings": {},<br> "helm_values": {},<br> "k8s_namespace": "kube-system",<br> "scrape_interval": "60s",<br> "scrape_timeout": "15s"<br>}</pre> | no | | <a name="input_ksm_config"></a> [ksm\_config](#input\_ksm\_config) | Kube State metrics configuration | <pre>object({<br> create_namespace = bool<br> k8s_namespace = string<br> helm_chart_name = string<br> helm_chart_version = string<br> helm_release_name = string<br> helm_repo_url = string<br> helm_settings = map(string)<br> helm_values = map(any)<br><br> scrape_interval = string<br> scrape_timeout = string<br> })</pre> | <pre>{<br> "create_namespace": true,<br> "helm_chart_name": "kube-state-metrics",<br> "helm_chart_version": "4.24.0",<br> "helm_release_name": "kube-state-metrics",<br> "helm_repo_url": "https://prometheus-community.github.io/helm-charts",<br> "helm_settings": {},<br> "helm_values": {},<br> "k8s_namespace": "kube-system",<br> "scrape_interval": "60s",<br> "scrape_timeout": "15s"<br>}</pre> | no |
| <a name="input_managed_prometheus_workspace_endpoint"></a> [managed\_prometheus\_workspace\_endpoint](#input\_managed\_prometheus\_workspace\_endpoint) | Amazon Managed Prometheus Workspace Endpoint | `string` | `""` | no | | <a name="input_managed_prometheus_workspace_endpoint"></a> [managed\_prometheus\_workspace\_endpoint](#input\_managed\_prometheus\_workspace\_endpoint) | Amazon Managed Prometheus Workspace Endpoint | `string` | `""` | no |
| <a name="input_managed_prometheus_workspace_id"></a> [managed\_prometheus\_workspace\_id](#input\_managed\_prometheus\_workspace\_id) | Amazon Managed Prometheus Workspace ID | `string` | `null` | no | | <a name="input_managed_prometheus_workspace_id"></a> [managed\_prometheus\_workspace\_id](#input\_managed\_prometheus\_workspace\_id) | Amazon Managed Prometheus Workspace ID | `string` | `null` | no |
| <a name="input_managed_prometheus_workspace_region"></a> [managed\_prometheus\_workspace\_region](#input\_managed\_prometheus\_workspace\_region) | Amazon Managed Prometheus Workspace's Region | `string` | `null` | no | | <a name="input_managed_prometheus_workspace_region"></a> [managed\_prometheus\_workspace\_region](#input\_managed\_prometheus\_workspace\_region) | Amazon Managed Prometheus Workspace's Region | `string` | `null` | no |
| <a name="input_ne_config"></a> [ne\_config](#input\_ne\_config) | Node exporter configuration | <pre>object({<br> create_namespace = bool<br> k8s_namespace = string<br> helm_chart_name = string<br> helm_chart_version = string<br> helm_release_name = string<br> helm_repo_url = string<br> helm_settings = map(string)<br> helm_values = map(any)<br><br> scrape_interval = string<br> scrape_timeout = string<br> })</pre> | <pre>{<br> "create_namespace": true,<br> "helm_chart_name": "prometheus-node-exporter",<br> "helm_chart_version": "2.0.3",<br> "helm_release_name": "prometheus-node-exporter",<br> "helm_repo_url": "https://prometheus-community.github.io/helm-charts",<br> "helm_settings": {},<br> "helm_values": {},<br> "k8s_namespace": "prometheus-node-exporter",<br> "scrape_interval": "60s",<br> "scrape_timeout": "60s"<br>}</pre> | no | | <a name="input_ne_config"></a> [ne\_config](#input\_ne\_config) | Node exporter configuration | <pre>object({<br> create_namespace = bool<br> k8s_namespace = string<br> helm_chart_name = string<br> helm_chart_version = string<br> helm_release_name = string<br> helm_repo_url = string<br> helm_settings = map(string)<br> helm_values = map(any)<br><br> scrape_interval = string<br> scrape_timeout = string<br> })</pre> | <pre>{<br> "create_namespace": true,<br> "helm_chart_name": "prometheus-node-exporter",<br> "helm_chart_version": "2.0.3",<br> "helm_release_name": "prometheus-node-exporter",<br> "helm_repo_url": "https://prometheus-community.github.io/helm-charts",<br> "helm_settings": {},<br> "helm_values": {},<br> "k8s_namespace": "prometheus-node-exporter",<br> "scrape_interval": "60s",<br> "scrape_timeout": "60s"<br>}</pre> | no |
| <a name="input_nginx_config"></a> [nginx\_config](#input\_nginx\_config) | Configuration object for NGINX monitoring | <pre>object({<br> enable_alerting_rules = bool<br> scrape_sample_limit = number<br> prometheus_metrics_endpoint = string<br> })</pre> | <pre>{<br> "enable_alerting_rules": true,<br> "prometheus_metrics_endpoint": "metrics",<br> "scrape_sample_limit": 1000<br>}</pre> | no |
| <a name="input_prometheus_config"></a> [prometheus\_config](#input\_prometheus\_config) | Controls default values such as scrape interval, timeouts and ports globally | <pre>object({<br> global_scrape_interval = string<br> global_scrape_timeout = string<br> })</pre> | <pre>{<br> "global_scrape_interval": "60s",<br> "global_scrape_timeout": "15s"<br>}</pre> | no | | <a name="input_prometheus_config"></a> [prometheus\_config](#input\_prometheus\_config) | Controls default values such as scrape interval, timeouts and ports globally | <pre>object({<br> global_scrape_interval = string<br> global_scrape_timeout = string<br> })</pre> | <pre>{<br> "global_scrape_interval": "60s",<br> "global_scrape_timeout": "15s"<br>}</pre> | no |
| <a name="input_tags"></a> [tags](#input\_tags) | Additional tags (e.g. `map('BusinessUnit`,`XYZ`) | `map(string)` | `{}` | no | | <a name="input_tags"></a> [tags](#input\_tags) | Additional tags (e.g. `map('BusinessUnit`,`XYZ`) | `map(string)` | `{}` | no |
| <a name="input_tracing_config"></a> [tracing\_config](#input\_tracing\_config) | Configuration object for traces collection to AWS X-Ray | <pre>object({<br> otlp_grpc_endpoint = string<br> otlp_http_endpoint = string<br> send_batch_size = number<br> timeout = string<br> })</pre> | <pre>{<br> "otlp_grpc_endpoint": "0.0.0.0:4317",<br> "otlp_http_endpoint": "0.0.0.0:4318",<br> "send_batch_size": 50,<br> "timeout": "30s"<br>}</pre> | no | | <a name="input_tracing_config"></a> [tracing\_config](#input\_tracing\_config) | Configuration object for traces collection to AWS X-Ray | <pre>object({<br> otlp_grpc_endpoint = string<br> otlp_http_endpoint = string<br> send_batch_size = number<br> timeout = string<br> })</pre> | <pre>{<br> "otlp_grpc_endpoint": "0.0.0.0:4317",<br> "otlp_http_endpoint": "0.0.0.0:4318",<br> "send_batch_size": 50,<br> "timeout": "30s"<br>}</pre> | no |
@@ -83,5 +91,7 @@ This module is inspired from the open source [kube-prometheus-stack](https://git
| Name | Description | | Name | Description |
|------|-------------| |------|-------------|
| <a name="output_eks_cluster_id"></a> [eks\_cluster\_id](#output\_eks\_cluster\_id) | EKS Cluster Id |
| <a name="output_eks_cluster_version"></a> [eks\_cluster\_version](#output\_eks\_cluster\_version) | EKS Cluster version |
| <a name="output_grafana_dashboard_urls"></a> [grafana\_dashboard\_urls](#output\_grafana\_dashboard\_urls) | URLs for dashboards created | | <a name="output_grafana_dashboard_urls"></a> [grafana\_dashboard\_urls](#output\_grafana\_dashboard\_urls) | URLs for dashboards created |
<!-- END OF PRE-COMMIT-TERRAFORM DOCS HOOK --> <!-- END OF PRE-COMMIT-TERRAFORM DOCS HOOK -->
@@ -14,6 +14,7 @@ locals {
eks_oidc_issuer_url = replace(data.aws_eks_cluster.eks_cluster.identity[0].oidc[0].issuer, "https://", "") eks_oidc_issuer_url = replace(data.aws_eks_cluster.eks_cluster.identity[0].oidc[0].issuer, "https://", "")
eks_cluster_endpoint = data.aws_eks_cluster.eks_cluster.endpoint eks_cluster_endpoint = data.aws_eks_cluster.eks_cluster.endpoint
eks_cluster_version = data.aws_eks_cluster.eks_cluster.version
context = { context = {
aws_caller_identity_account_id = data.aws_caller_identity.current.account_id aws_caller_identity_account_id = data.aws_caller_identity.current.account_id
@@ -1,3 +1,12 @@
module "operator" {
source = "./add-ons/adot-operator"
count = var.enable_amazon_eks_adot ? 1 : 0
enable_cert_manager = var.enable_cert_manager
kubernetes_version = local.eks_cluster_version
addon_context = local.context
}
resource "helm_release" "kube_state_metrics" { resource "helm_release" "kube_state_metrics" {
count = var.enable_kube_state_metrics ? 1 : 0 count = var.enable_kube_state_metrics ? 1 : 0
chart = var.ksm_config.helm_chart_name chart = var.ksm_config.helm_chart_name
@@ -104,6 +113,26 @@ module "helm_addon" {
{ {
name = "customMetricsDroppedSeriesPrefixes" name = "customMetricsDroppedSeriesPrefixes"
value = format("(%s.*)$", join(".*|", var.custom_metrics_config.dropped_series_prefixes)) value = format("(%s.*)$", join(".*|", var.custom_metrics_config.dropped_series_prefixes))
},
{
name = "enable_java"
value = var.enable_java
},
{
name = "javaScrapeSampleLimit"
value = var.java_config.scrape_sample_limit
},
{
name = "enable_nginx"
value = var.enable_nginx
},
{
name = "nginxScrapeSampleLimit"
value = var.nginx_config.scrape_sample_limit
},
{
name = "nginxPrometheusMetricsEndpoint"
value = var.nginx_config.prometheus_metrics_endpoint
} }
] ]
@@ -119,4 +148,24 @@ module "helm_addon" {
} }
addon_context = local.context addon_context = local.context
depends_on = [module.operator]
}
module "java_monitoring" {
source = "./patterns/java"
count = var.enable_java ? 1 : 0
managed_prometheus_workspace_id = var.managed_prometheus_workspace_id
enable_alerting_rules = var.java_config.enable_alerting_rules
dashboards_folder_id = var.dashboards_folder_id
}
module "nginx_monitoring" {
source = "./patterns/nginx"
count = var.enable_nginx ? 1 : 0
managed_prometheus_workspace_id = var.managed_prometheus_workspace_id
enable_alerting_rules = var.nginx_config.enable_alerting_rules
dashboards_folder_id = var.dashboards_folder_id
} }
@@ -1703,6 +1703,72 @@ spec:
- source_labels: [ __name__ ] - source_labels: [ __name__ ]
regex: '{{ .Values.customMetricsDroppedSeriesPrefixes }}' regex: '{{ .Values.customMetricsDroppedSeriesPrefixes }}'
action: drop action: drop
{{ if .Values.enableJava }}
- job_name: 'kubernetes-java-jmx'
sample_limit: {{ .Values.javaScrapeSampleLimit }}
kubernetes_sd_configs:
- role: pod
relabel_configs:
- source_labels: [ __address__ ]
action: keep
regex: '.*:9404$'
- action: labelmap
regex: __meta_kubernetes_pod_label_(.+)
- action: replace
source_labels: [ __meta_kubernetes_namespace ]
target_label: Namespace
- source_labels: [ __meta_kubernetes_pod_name ]
action: replace
target_label: pod_name
- action: replace
source_labels: [ __meta_kubernetes_pod_container_name ]
target_label: container_name
- action: replace
source_labels: [ __meta_kubernetes_pod_controller_kind ]
target_label: pod_controller_kind
- action: replace
source_labels: [ __meta_kubernetes_pod_phase ]
target_label: pod_controller_phase
metric_relabel_configs:
- source_labels: [ __name__ ]
regex: 'jvm_gc_collection_seconds.*'
action: drop
{{ end }}
{{ if .Values.enableNginx }}
- job_name: 'kubernetes-nginx'
sample_limit: {{ .Values.nginxScrapeSampleLimit }}
metrics_path: /{{ .Values.nginxPrometheusMetricsEndpoint }}
kubernetes_sd_configs:
- role: pod
relabel_configs:
- source_labels: [ __address__ ]
action: keep
regex: '.*:10254$'
- source_labels: [__meta_kubernetes_pod_container_name]
target_label: container
action: replace
- source_labels: [__meta_kubernetes_pod_node_name]
target_label: host
action: replace
- source_labels: [__meta_kubernetes_namespace]
target_label: namespace
action: replace
metric_relabel_configs:
- source_labels: [__name__]
regex: 'go_memstats.*'
action: drop
- source_labels: [__name__]
regex: 'go_gc.*'
action: drop
- source_labels: [__name__]
regex: 'go_threads'
action: drop
- regex: exported_host
action: labeldrop
{{ end }}
exporters: exporters:
{{ if .Values.enableTracing }} {{ if .Values.enableTracing }}
awsxray: awsxray:
@@ -14,3 +14,10 @@ tracingSendBatchSize: ${tracing_send_batch_size}
enableCustomMetrics: ${enable_custom_metrics} enableCustomMetrics: ${enable_custom_metrics}
customMetricsPorts: ${custom_metrics_ports} customMetricsPorts: ${custom_metrics_ports}
customMetricsDroppedSeriesPrefixes: ${custom_metrics_dropped_series_prefixes} customMetricsDroppedSeriesPrefixes: ${custom_metrics_dropped_series_prefixes}
enableJava: ${enable_java}
javaScrapeSampleLimit: ${java_scrape_sample_limit}
enableNginx: ${enable_nginx}
nginxScrapeSampleLimit: ${nginx_scrape_sample_limit}
nginxPrometheusMetricsEndpoint: ${nginx_prometheus_metrics_endpoint}
+21
View File
@@ -0,0 +1,21 @@
output "grafana_dashboard_urls" {
value = [concat(
grafana_dashboard.workloads[*].url,
grafana_dashboard.nodes[*].url,
grafana_dashboard.nsworkload[*].url,
grafana_dashboard.kubelet[*].url,
grafana_dashboard.cluster[*].url,
flatten(module.java_monitoring[*].grafana_dashboard_urls),
flatten(module.nginx_monitoring[*].grafana_dashboard_urls),
)]
description = "URLs for dashboards created"
}
output "eks_cluster_version" {
description = "EKS Cluster version"
value = data.aws_eks_cluster.eks_cluster.version
}
output "eks_cluster_id" {
description = "EKS Cluster Id"
value = var.eks_cluster_id
}
@@ -0,0 +1,52 @@
# Java patterns module
Provides monitoring for Java based workloads with the following resources:
- AWS Managed Grafana Dashboard and data source
- Alerts and recording rules with AWS Managed Service for Prometheus
<!-- BEGINNING OF PRE-COMMIT-TERRAFORM DOCS HOOK -->
## Requirements
| Name | Version |
|------|---------|
| <a name="requirement_terraform"></a> [terraform](#requirement\_terraform) | >= 1.1.0 |
| <a name="requirement_aws"></a> [aws](#requirement\_aws) | >= 4.0.0 |
| <a name="requirement_grafana"></a> [grafana](#requirement\_grafana) | >= 1.25.0 |
| <a name="requirement_helm"></a> [helm](#requirement\_helm) | >= 2.4.1 |
| <a name="requirement_kubectl"></a> [kubectl](#requirement\_kubectl) | >= 1.14 |
| <a name="requirement_kubernetes"></a> [kubernetes](#requirement\_kubernetes) | >= 2.10 |
## Providers
| Name | Version |
|------|---------|
| <a name="provider_aws"></a> [aws](#provider\_aws) | >= 4.0.0 |
| <a name="provider_grafana"></a> [grafana](#provider\_grafana) | >= 1.25.0 |
## Modules
No modules.
## Resources
| Name | Type |
|------|------|
| [aws_prometheus_rule_group_namespace.alerting_rules](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/resources/prometheus_rule_group_namespace) | resource |
| [aws_prometheus_rule_group_namespace.recording_rules](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/resources/prometheus_rule_group_namespace) | resource |
| [grafana_dashboard.this](https://registry.terraform.io/providers/grafana/grafana/latest/docs/resources/dashboard) | resource |
## Inputs
| Name | Description | Type | Default | Required |
|------|-------------|------|---------|:--------:|
| <a name="input_dashboards_folder_id"></a> [dashboards\_folder\_id](#input\_dashboards\_folder\_id) | Grafana folder ID for automatic dashboards | `string` | n/a | yes |
| <a name="input_enable_alerting_rules"></a> [enable\_alerting\_rules](#input\_enable\_alerting\_rules) | Enables or disables Managed Prometheus alerting rules | `bool` | `true` | no |
| <a name="input_managed_prometheus_workspace_id"></a> [managed\_prometheus\_workspace\_id](#input\_managed\_prometheus\_workspace\_id) | Amazon Managed Prometheus Workspace ID | `string` | `null` | no |
## Outputs
| Name | Description |
|------|-------------|
| <a name="output_grafana_dashboard_urls"></a> [grafana\_dashboard\_urls](#output\_grafana\_dashboard\_urls) | URLs for dashboards created |
<!-- END OF PRE-COMMIT-TERRAFORM DOCS HOOK -->
@@ -0,0 +1,36 @@
resource "aws_prometheus_rule_group_namespace" "recording_rules" {
name = "accelerator-java-rules"
workspace_id = var.managed_prometheus_workspace_id
data = <<EOF
groups:
- name: default-metric
rules:
- record: metric:recording_rule
expr: avg(rate(container_cpu_usage_seconds_total[5m]))
EOF
}
resource "aws_prometheus_rule_group_namespace" "alerting_rules" {
count = var.enable_alerting_rules ? 1 : 0
name = "accelerator-java-alerting"
workspace_id = var.managed_prometheus_workspace_id
data = <<EOF
groups:
- name: default-alert
rules:
- alert: metric:alerting_rule
expr: jvm_memory_bytes_used{job="java", area="heap"} / jvm_memory_bytes_max * 100 > 80
for: 1m
labels:
severity: warning
annotations:
summary: "JVM heap warning"
description: "JVM heap of instance `{{$labels.instance}}` from application `{{$labels.application}}` is above 80% for one minute. (current=`{{$value}}%`)"
EOF
}
resource "grafana_dashboard" "this" {
folder = var.dashboards_folder_id
config_json = file("${path.module}/dashboards/default.json")
}
@@ -0,0 +1,16 @@
variable "enable_alerting_rules" {
description = "Enables or disables Managed Prometheus alerting rules"
type = bool
default = true
}
variable "managed_prometheus_workspace_id" {
description = "Amazon Managed Prometheus Workspace ID"
type = string
default = null
}
variable "dashboards_folder_id" {
description = "Grafana folder ID for automatic dashboards"
type = string
}
@@ -26,9 +26,7 @@ It provides the following resources:
## Modules ## Modules
| Name | Source | Version | No modules.
|------|--------|---------|
| <a name="module_helm_addon"></a> [helm\_addon](#module\_helm\_addon) | github.com/aws-ia/terraform-aws-eks-blueprints//modules/kubernetes-addons/helm-addon | v4.13.1 |
## Resources ## Resources
@@ -36,27 +34,14 @@ It provides the following resources:
|------|------| |------|------|
| [aws_prometheus_rule_group_namespace.alerting_rules](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/resources/prometheus_rule_group_namespace) | resource | | [aws_prometheus_rule_group_namespace.alerting_rules](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/resources/prometheus_rule_group_namespace) | resource |
| [grafana_dashboard.workloads](https://registry.terraform.io/providers/grafana/grafana/latest/docs/resources/dashboard) | resource | | [grafana_dashboard.workloads](https://registry.terraform.io/providers/grafana/grafana/latest/docs/resources/dashboard) | resource |
| [aws_caller_identity.current](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/data-sources/caller_identity) | data source |
| [aws_eks_cluster.eks_cluster](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/data-sources/eks_cluster) | data source |
| [aws_partition.current](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/data-sources/partition) | data source |
| [aws_region.current](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/data-sources/region) | data source |
## Inputs ## Inputs
| Name | Description | Type | Default | Required | | Name | Description | Type | Default | Required |
|------|-------------|------|---------|:--------:| |------|-------------|------|---------|:--------:|
| <a name="input_config"></a> [config](#input\_config) | Helm Config for Prometheus | `any` | `{}` | no |
| <a name="input_dashboards_folder_id"></a> [dashboards\_folder\_id](#input\_dashboards\_folder\_id) | Grafana folder ID for automatic dashboards | `string` | n/a | yes | | <a name="input_dashboards_folder_id"></a> [dashboards\_folder\_id](#input\_dashboards\_folder\_id) | Grafana folder ID for automatic dashboards | `string` | n/a | yes |
| <a name="input_eks_cluster_id"></a> [eks\_cluster\_id](#input\_eks\_cluster\_id) | EKS Cluster Id | `string` | n/a | yes |
| <a name="input_enable_alerting_rules"></a> [enable\_alerting\_rules](#input\_enable\_alerting\_rules) | Enables or disables Managed Prometheus alerting rules | `bool` | `true` | no | | <a name="input_enable_alerting_rules"></a> [enable\_alerting\_rules](#input\_enable\_alerting\_rules) | Enables or disables Managed Prometheus alerting rules | `bool` | `true` | no |
| <a name="input_enable_dashboards"></a> [enable\_dashboards](#input\_enable\_dashboards) | Enables or disables curated dashboards | `bool` | `true` | no |
| <a name="input_helm_config"></a> [helm\_config](#input\_helm\_config) | Helm Config for Prometheus | `any` | `{}` | no |
| <a name="input_irsa_iam_permissions_boundary"></a> [irsa\_iam\_permissions\_boundary](#input\_irsa\_iam\_permissions\_boundary) | IAM permissions boundary for IRSA roles | `string` | `null` | no |
| <a name="input_irsa_iam_role_path"></a> [irsa\_iam\_role\_path](#input\_irsa\_iam\_role\_path) | IAM role path for IRSA roles | `string` | `"/"` | no |
| <a name="input_managed_prometheus_workspace_endpoint"></a> [managed\_prometheus\_workspace\_endpoint](#input\_managed\_prometheus\_workspace\_endpoint) | Amazon Managed Prometheus Workspace Endpoint | `string` | `""` | no |
| <a name="input_managed_prometheus_workspace_id"></a> [managed\_prometheus\_workspace\_id](#input\_managed\_prometheus\_workspace\_id) | Amazon Managed Prometheus Workspace ID | `string` | `null` | no | | <a name="input_managed_prometheus_workspace_id"></a> [managed\_prometheus\_workspace\_id](#input\_managed\_prometheus\_workspace\_id) | Amazon Managed Prometheus Workspace ID | `string` | `null` | no |
| <a name="input_managed_prometheus_workspace_region"></a> [managed\_prometheus\_workspace\_region](#input\_managed\_prometheus\_workspace\_region) | Amazon Managed Prometheus Workspace's Region | `string` | `null` | no |
| <a name="input_tags"></a> [tags](#input\_tags) | Additional tags (e.g. `map('BusinessUnit`,`XYZ`) | `map(string)` | `{}` | no |
## Outputs ## Outputs
@@ -1,7 +1,3 @@
################################################################################################################################################
# Alerting rules ###############################################################################################################################
################################################################################################################################################
resource "aws_prometheus_rule_group_namespace" "alerting_rules" { resource "aws_prometheus_rule_group_namespace" "alerting_rules" {
count = var.enable_alerting_rules ? 1 : 0 count = var.enable_alerting_rules ? 1 : 0
@@ -41,3 +37,8 @@ groups:
description: "Nginx p99 latency is higher than 3 seconds\n VALUE = {{ $value }}\n LABELS = {{ $labels }}" description: "Nginx p99 latency is higher than 3 seconds\n VALUE = {{ $value }}\n LABELS = {{ $labels }}"
EOF EOF
} }
resource "grafana_dashboard" "workloads" {
folder = var.dashboards_folder_id
config_json = file("${path.module}/dashboards/nginx.json")
}
@@ -0,0 +1,16 @@
variable "managed_prometheus_workspace_id" {
description = "Amazon Managed Prometheus Workspace ID"
type = string
default = null
}
variable "dashboards_folder_id" {
type = string
description = "Grafana folder ID for automatic dashboards"
}
variable "enable_alerting_rules" {
type = bool
default = true
description = "Enables or disables Managed Prometheus alerting rules"
}
@@ -5,7 +5,6 @@
################################################################################################################################################ ################################################################################################################################################
resource "aws_prometheus_rule_group_namespace" "recording_rules" { resource "aws_prometheus_rule_group_namespace" "recording_rules" {
count = var.enable_recording_rules ? 1 : 0
name = "accelerator-infra-rules" name = "accelerator-infra-rules"
workspace_id = var.managed_prometheus_workspace_id workspace_id = var.managed_prometheus_workspace_id
data = <<EOF data = <<EOF
@@ -3,6 +3,18 @@ variable "eks_cluster_id" {
type = string type = string
} }
variable "enable_amazon_eks_adot" {
description = "Enables the ADOT Operator on the EKS Cluster"
type = bool
default = true
}
variable "enable_cert_manager" {
description = "Allow reusing an existing installation of cert-manager"
type = bool
default = true
}
variable "helm_config" { variable "helm_config" {
description = "Helm Config for Prometheus" description = "Helm Config for Prometheus"
type = any type = any
@@ -44,12 +56,6 @@ variable "dashboards_folder_id" {
type = string type = string
} }
variable "enable_recording_rules" {
description = "Enables or disables Managed Prometheus recording rules. Disabling this might affect some data in the dashboards"
type = bool
default = true
}
variable "enable_alerting_rules" { variable "enable_alerting_rules" {
description = "Enables or disables Managed Prometheus alerting rules" description = "Enables or disables Managed Prometheus alerting rules"
type = bool type = bool
@@ -201,3 +207,43 @@ variable "custom_metrics_config" {
dropped_series_prefixes = ["unspecified"] dropped_series_prefixes = ["unspecified"]
} }
} }
variable "enable_java" {
description = "Enable Java workloads monitoring, alerting and default dashboards"
type = bool
default = false
}
variable "java_config" {
description = "Configuration object for Java/JMX monitoring"
type = object({
enable_alerting_rules = bool
scrape_sample_limit = number
})
default = {
enable_alerting_rules = true
scrape_sample_limit = 1000
}
}
variable "enable_nginx" {
description = "Enable NGINX workloads monitoring, alerting and default dashboards"
type = bool
default = false
}
variable "nginx_config" {
description = "Configuration object for NGINX monitoring"
type = object({
enable_alerting_rules = bool
scrape_sample_limit = number
prometheus_metrics_endpoint = string
})
default = {
enable_alerting_rules = true
scrape_sample_limit = 1000
prometheus_metrics_endpoint = "metrics"
}
}
@@ -9,9 +9,11 @@ locals {
} }
resource "grafana_data_source" "cloudwatch" { resource "grafana_data_source" "cloudwatch" {
type = "cloudwatch" type = "cloudwatch"
name = local.name name = local.name
is_default = true
# Giving priority to Managed Prometheus datasources
is_default = false
json_data { json_data {
default_region = var.aws_region default_region = var.aws_region
sigv4_auth = true sigv4_auth = true
@@ -26,7 +28,7 @@ resource "grafana_dashboard" "this" {
} }
module "billing" { module "billing" {
source = "../../workloads/managed-prometheus-monitoring/billing" source = "./billing"
providers = { providers = {
aws = aws.billing_region aws = aws.billing_region
} }
-10
View File
@@ -1,10 +0,0 @@
output "grafana_dashboard_urls" {
value = [concat(
grafana_dashboard.workloads[*].url,
grafana_dashboard.nodes[*].url,
grafana_dashboard.nsworkload[*].url,
grafana_dashboard.kubelet[*].url,
grafana_dashboard.cluster[*].url,
)]
description = "URLs for dashboards created"
}
-68
View File
@@ -1,68 +0,0 @@
# Java based workloads monitoring
This module provides monitoring for Java based workloads with the following resources:
- AWS Distro For OpenTelemetry Operator and Collector
- AWS Managed Grafana Dashboard and data source
- Alerts and recording rules with AWS Managed Service for Prometheus
<!-- BEGINNING OF PRE-COMMIT-TERRAFORM DOCS HOOK -->
## Requirements
| Name | Version |
|------|---------|
| <a name="requirement_terraform"></a> [terraform](#requirement\_terraform) | >= 1.1.0 |
| <a name="requirement_aws"></a> [aws](#requirement\_aws) | >= 4.0.0 |
| <a name="requirement_grafana"></a> [grafana](#requirement\_grafana) | >= 1.25.0 |
| <a name="requirement_helm"></a> [helm](#requirement\_helm) | >= 2.4.1 |
| <a name="requirement_kubectl"></a> [kubectl](#requirement\_kubectl) | >= 1.14 |
| <a name="requirement_kubernetes"></a> [kubernetes](#requirement\_kubernetes) | >= 2.10 |
## Providers
| Name | Version |
|------|---------|
| <a name="provider_aws"></a> [aws](#provider\_aws) | >= 4.0.0 |
| <a name="provider_grafana"></a> [grafana](#provider\_grafana) | >= 1.25.0 |
## Modules
| Name | Source | Version |
|------|--------|---------|
| <a name="module_helm_addon"></a> [helm\_addon](#module\_helm\_addon) | github.com/aws-ia/terraform-aws-eks-blueprints//modules/kubernetes-addons/helm-addon | v4.13.1 |
## Resources
| Name | Type |
|------|------|
| [aws_prometheus_rule_group_namespace.alerting_rules](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/resources/prometheus_rule_group_namespace) | resource |
| [aws_prometheus_rule_group_namespace.recording_rules](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/resources/prometheus_rule_group_namespace) | resource |
| [grafana_dashboard.this](https://registry.terraform.io/providers/grafana/grafana/latest/docs/resources/dashboard) | resource |
| [aws_caller_identity.current](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/data-sources/caller_identity) | data source |
| [aws_eks_cluster.eks_cluster](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/data-sources/eks_cluster) | data source |
| [aws_partition.current](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/data-sources/partition) | data source |
| [aws_region.current](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/data-sources/region) | data source |
## Inputs
| Name | Description | Type | Default | Required |
|------|-------------|------|---------|:--------:|
| <a name="input_dashboards_folder_id"></a> [dashboards\_folder\_id](#input\_dashboards\_folder\_id) | Grafana folder ID for automatic dashboards | `string` | n/a | yes |
| <a name="input_eks_cluster_id"></a> [eks\_cluster\_id](#input\_eks\_cluster\_id) | EKS Cluster Id | `string` | n/a | yes |
| <a name="input_enable_alerting_rules"></a> [enable\_alerting\_rules](#input\_enable\_alerting\_rules) | Enables or disables Managed Prometheus alerting rules | `bool` | `true` | no |
| <a name="input_enable_recording_rules"></a> [enable\_recording\_rules](#input\_enable\_recording\_rules) | Enables or disables Managed Prometheus recording rules. Disabling this might affect some data in the dashboards | `bool` | `true` | no |
| <a name="input_helm_config"></a> [helm\_config](#input\_helm\_config) | Helm Config for Prometheus | `any` | `{}` | no |
| <a name="input_irsa_iam_permissions_boundary"></a> [irsa\_iam\_permissions\_boundary](#input\_irsa\_iam\_permissions\_boundary) | IAM permissions boundary for IRSA roles | `string` | `null` | no |
| <a name="input_irsa_iam_role_path"></a> [irsa\_iam\_role\_path](#input\_irsa\_iam\_role\_path) | IAM role path for IRSA roles | `string` | `"/"` | no |
| <a name="input_managed_prometheus_workspace_endpoint"></a> [managed\_prometheus\_workspace\_endpoint](#input\_managed\_prometheus\_workspace\_endpoint) | Amazon Managed Prometheus Workspace Endpoint | `string` | `""` | no |
| <a name="input_managed_prometheus_workspace_id"></a> [managed\_prometheus\_workspace\_id](#input\_managed\_prometheus\_workspace\_id) | Amazon Managed Prometheus Workspace ID | `string` | `null` | no |
| <a name="input_managed_prometheus_workspace_region"></a> [managed\_prometheus\_workspace\_region](#input\_managed\_prometheus\_workspace\_region) | Amazon Managed Prometheus Workspace's Region | `string` | `null` | no |
| <a name="input_prometheus_config"></a> [prometheus\_config](#input\_prometheus\_config) | Controls default values such as scrape interval, timeouts and ports globally | <pre>object({<br> global_scrape_interval = string<br> global_scrape_timeout = string<br> scrape_sample_limit = number<br> })</pre> | <pre>{<br> "global_scrape_interval": "60s",<br> "global_scrape_timeout": "15s",<br> "scrape_sample_limit": 1000<br>}</pre> | no |
| <a name="input_tags"></a> [tags](#input\_tags) | Additional tags (e.g. `map('BusinessUnit`,`XYZ`) | `map(string)` | `{}` | no |
## Outputs
| Name | Description |
|------|-------------|
| <a name="output_grafana_dashboard_urls"></a> [grafana\_dashboard\_urls](#output\_grafana\_dashboard\_urls) | URLs for dashboards created |
<!-- END OF PRE-COMMIT-TERRAFORM DOCS HOOK -->
-31
View File
@@ -1,31 +0,0 @@
data "aws_partition" "current" {}
data "aws_caller_identity" "current" {}
data "aws_region" "current" {}
data "aws_eks_cluster" "eks_cluster" {
name = var.eks_cluster_id
}
locals {
name = "adot-collector-java"
namespace = try(var.helm_config.namespace, local.name)
eks_oidc_issuer_url = replace(data.aws_eks_cluster.eks_cluster.identity[0].oidc[0].issuer, "https://", "")
eks_cluster_endpoint = data.aws_eks_cluster.eks_cluster.endpoint
context = {
aws_caller_identity_account_id = data.aws_caller_identity.current.account_id
aws_caller_identity_arn = data.aws_caller_identity.current.arn
aws_eks_cluster_endpoint = local.eks_cluster_endpoint
aws_partition_id = data.aws_partition.current.partition
aws_region_name = data.aws_region.current.name
eks_cluster_id = var.eks_cluster_id
eks_oidc_issuer_url = local.eks_oidc_issuer_url
eks_oidc_provider_arn = "arn:${data.aws_partition.current.partition}:iam::${data.aws_caller_identity.current.account_id}:oidc-provider/${local.eks_oidc_issuer_url}"
tags = var.tags
irsa_iam_role_path = var.irsa_iam_role_path
irsa_iam_permissions_boundary = var.irsa_iam_permissions_boundary
}
}
-95
View File
@@ -1,95 +0,0 @@
# deploys collector
module "helm_addon" {
source = "github.com/aws-ia/terraform-aws-eks-blueprints//modules/kubernetes-addons/helm-addon?ref=v4.13.1"
helm_config = merge(
{
name = local.name
chart = "${path.module}/otel-config"
version = "0.2.0"
namespace = local.namespace
description = "ADOT helm Chart deployment configuration"
},
var.helm_config
)
set_values = [
{
name = "ampurl"
value = "${var.managed_prometheus_workspace_endpoint}api/v1/remote_write"
},
{
name = "region"
value = var.managed_prometheus_workspace_region
},
{
name = "ekscluster"
value = local.context.eks_cluster_id
},
{
name = "accountId"
value = local.context.aws_caller_identity_account_id
},
{
name = "globalScrapeInterval"
value = var.prometheus_config.global_scrape_interval
},
{
name = "globalScrapeTimeout"
value = var.prometheus_config.global_scrape_timeout
},
{
name = "scrapeSampleLimit"
value = var.prometheus_config.scrape_sample_limit
}
]
irsa_config = {
create_kubernetes_namespace = try(var.helm_config["create_namespace"], true)
kubernetes_namespace = local.namespace
create_kubernetes_service_account = true
kubernetes_service_account = try(var.helm_config.service_account, local.name)
irsa_iam_policies = ["arn:${data.aws_partition.current.partition}:iam::aws:policy/AmazonPrometheusRemoteWriteAccess"]
}
addon_context = local.context
}
resource "aws_prometheus_rule_group_namespace" "recording_rules" {
count = var.enable_recording_rules ? 1 : 0
name = "accelerator-java-rules"
workspace_id = var.managed_prometheus_workspace_id
data = <<EOF
groups:
- name: default-metric
rules:
- record: metric:recording_rule
expr: avg(rate(container_cpu_usage_seconds_total[5m]))
EOF
}
resource "aws_prometheus_rule_group_namespace" "alerting_rules" {
count = var.enable_alerting_rules ? 1 : 0
name = "accelerator-java-alerting"
workspace_id = var.managed_prometheus_workspace_id
data = <<EOF
groups:
- name: default-alert
rules:
- alert: metric:alerting_rule
expr: jvm_memory_bytes_used{job="java", area="heap"} / jvm_memory_bytes_max * 100 > 80
for: 1m
labels:
severity: warning
annotations:
summary: "JVM heap warning"
description: "JVM heap of instance `{{$labels.instance}}` from application `{{$labels.application}}` is above 80% for one minute. (current=`{{$value}}%`)"
EOF
}
resource "grafana_dashboard" "this" {
folder = var.dashboards_folder_id
config_json = file("${path.module}/dashboards/default.json")
}
@@ -1,6 +0,0 @@
apiVersion: v2
name: opentelemetry
description: A Helm chart to install otel operator
type: application
version: 0.2.0
appVersion: v0.1.0
@@ -1,29 +0,0 @@
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: otel-prometheus-role
rules:
- apiGroups:
- ""
resources:
- nodes
- nodes/proxy
- services
- endpoints
- pods
verbs:
- get
- list
- watch
- apiGroups:
- extensions
resources:
- ingresses
verbs:
- get
- list
- watch
- nonResourceURLs:
- /metrics
verbs:
- get
@@ -1,12 +0,0 @@
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
name: otel-prometheus-role-binding
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: ClusterRole
name: otel-prometheus-role
subjects:
- kind: ServiceAccount
name: adot-collector-java
namespace: adot-collector-java
@@ -1,71 +0,0 @@
apiVersion: opentelemetry.io/v1alpha1
kind: OpenTelemetryCollector
metadata:
name: adot
spec:
image: public.ecr.aws/aws-observability/aws-otel-collector:v0.22.0
mode: deployment
serviceAccount: adot-collector-java
config: |
receivers:
prometheus:
config:
global:
scrape_interval: {{ .Values.globalScrapeInterval }}
scrape_timeout: {{ .Values.globalScrapeTimeout }}
external_labels:
cluster: {{ .Values.ekscluster }}
account_id: {{ .Values.accountId }}
region: {{ .Values.region }}
scrape_configs:
- job_name: 'kubernetes-pod-jmx'
sample_limit: {{ .Values.scrapeSampleLimit }}
kubernetes_sd_configs:
- role: pod
relabel_configs:
- source_labels: [ __address__ ]
action: keep
regex: '.*:9404$'
- action: labelmap
regex: __meta_kubernetes_pod_label_(.+)
- action: replace
source_labels: [ __meta_kubernetes_namespace ]
target_label: Namespace
- source_labels: [ __meta_kubernetes_pod_name ]
action: replace
target_label: pod_name
- action: replace
source_labels: [ __meta_kubernetes_pod_container_name ]
target_label: container_name
- action: replace
source_labels: [ __meta_kubernetes_pod_controller_kind ]
target_label: pod_controller_kind
- action: replace
source_labels: [ __meta_kubernetes_pod_phase ]
target_label: pod_controller_phase
metric_relabel_configs:
- source_labels: [ __name__ ]
regex: 'jvm_gc_collection_seconds.*'
action: drop
exporters:
prometheusremotewrite:
endpoint: {{ .Values.ampurl }}
auth:
authenticator: sigv4auth
logging:
loglevel: info
extensions:
sigv4auth:
region: {{ .Values.region }}
service: "aps"
health_check:
pprof:
endpoint: :1888
zpages:
endpoint: :55679
service:
extensions: [pprof, zpages, health_check, sigv4auth]
pipelines:
metrics:
receivers: [prometheus]
exporters: [logging, prometheusremotewrite]
@@ -1,6 +0,0 @@
ampurl: ${amp_url}
region: ${region}
prometheusMetricsEndpoint: ${prometheus_metrics_endpoint}
globalScrapeInterval: ${scrape_interval}
globalScrapeTimeout: ${scrape_timeout}
scrapeSampleLimit: ${scrape_sample_limit}
-79
View File
@@ -1,79 +0,0 @@
variable "eks_cluster_id" {
description = "EKS Cluster Id"
type = string
}
variable "irsa_iam_role_path" {
description = "IAM role path for IRSA roles"
type = string
default = "/"
}
variable "irsa_iam_permissions_boundary" {
description = "IAM permissions boundary for IRSA roles"
type = string
default = null
}
variable "enable_recording_rules" {
description = "Enables or disables Managed Prometheus recording rules. Disabling this might affect some data in the dashboards"
type = bool
default = true
}
variable "enable_alerting_rules" {
description = "Enables or disables Managed Prometheus alerting rules"
type = bool
default = true
}
variable "managed_prometheus_workspace_endpoint" {
description = "Amazon Managed Prometheus Workspace Endpoint"
type = string
default = ""
}
variable "managed_prometheus_workspace_id" {
description = "Amazon Managed Prometheus Workspace ID"
type = string
default = null
}
variable "managed_prometheus_workspace_region" {
description = "Amazon Managed Prometheus Workspace's Region"
type = string
default = null
}
variable "helm_config" {
description = "Helm Config for Prometheus"
type = any
default = {}
}
variable "dashboards_folder_id" {
description = "Grafana folder ID for automatic dashboards"
type = string
}
variable "prometheus_config" {
description = "Controls default values such as scrape interval, timeouts and ports globally"
type = object({
global_scrape_interval = string
global_scrape_timeout = string
scrape_sample_limit = number
})
default = {
global_scrape_interval = "60s"
global_scrape_timeout = "15s"
scrape_sample_limit = 1000
}
nullable = false
}
variable "tags" {
description = "Additional tags (e.g. `map('BusinessUnit`,`XYZ`)"
type = map(string)
default = {}
}
-6
View File
@@ -1,6 +0,0 @@
resource "grafana_dashboard" "workloads" {
count = var.enable_dashboards ? 1 : 0
folder = var.dashboards_folder_id
config_json = file("${path.module}/dashboards/nginx.json")
}
-31
View File
@@ -1,31 +0,0 @@
data "aws_partition" "current" {}
data "aws_caller_identity" "current" {}
data "aws_region" "current" {}
data "aws_eks_cluster" "eks_cluster" {
name = var.eks_cluster_id
}
locals {
name = "adot-collector-nginx"
namespace = try(var.config.helm_config.namespace, local.name)
eks_oidc_issuer_url = replace(data.aws_eks_cluster.eks_cluster.identity[0].oidc[0].issuer, "https://", "")
eks_cluster_endpoint = data.aws_eks_cluster.eks_cluster.endpoint
context = {
aws_caller_identity_account_id = data.aws_caller_identity.current.account_id
aws_caller_identity_arn = data.aws_caller_identity.current.arn
aws_eks_cluster_endpoint = local.eks_cluster_endpoint
aws_partition_id = data.aws_partition.current.partition
aws_region_name = data.aws_region.current.name
eks_cluster_id = var.eks_cluster_id
eks_oidc_issuer_url = local.eks_oidc_issuer_url
eks_oidc_provider_arn = "arn:${data.aws_partition.current.partition}:iam::${data.aws_caller_identity.current.account_id}:oidc-provider/${local.eks_oidc_issuer_url}"
tags = var.tags
irsa_iam_role_path = var.irsa_iam_role_path
irsa_iam_permissions_boundary = var.irsa_iam_permissions_boundary
}
}
-55
View File
@@ -1,55 +0,0 @@
module "helm_addon" {
source = "github.com/aws-ia/terraform-aws-eks-blueprints//modules/kubernetes-addons/helm-addon?ref=v4.13.1"
helm_config = merge(
{
name = local.name
chart = "${path.module}/otel-config"
version = "0.2.0"
namespace = local.namespace
description = "ADOT helm Chart deployment configuration"
},
var.helm_config
)
set_values = [
{
name = "ampurl"
value = "${var.managed_prometheus_workspace_endpoint}api/v1/remote_write"
},
{
name = "region"
value = var.managed_prometheus_workspace_region
},
{
name = "prometheusMetricsEndpoint"
value = "metrics"
},
{
name = "prometheusMetricsPort"
value = 8888
},
{
name = "scrapeInterval"
value = "15s"
},
{
name = "scrapeTimeout"
value = "10s"
},
{
name = "scrapeSampleLimit"
value = 1000
}
]
irsa_config = {
create_kubernetes_namespace = try(var.helm_config["create_namespace"], true)
kubernetes_namespace = local.namespace
create_kubernetes_service_account = true
kubernetes_service_account = try(var.helm_config.service_account, local.name)
irsa_iam_policies = ["arn:${data.aws_partition.current.partition}:iam::aws:policy/AmazonPrometheusRemoteWriteAccess"]
}
addon_context = local.context
}
@@ -1,6 +0,0 @@
apiVersion: v2
name: opentelemetry
description: A Helm chart to install otel operator
type: application
version: 0.2.0
appVersion: v0.1.0
@@ -1,29 +0,0 @@
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: otel-prometheus-role
rules:
- apiGroups:
- ""
resources:
- nodes
- nodes/proxy
- services
- endpoints
- pods
verbs:
- get
- list
- watch
- apiGroups:
- extensions
resources:
- ingresses
verbs:
- get
- list
- watch
- nonResourceURLs:
- /metrics
verbs:
- get
@@ -1,12 +0,0 @@
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
name: otel-prometheus-role-binding
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: ClusterRole
name: otel-prometheus-role
subjects:
- kind: ServiceAccount
name: adot-collector-nginx
namespace: adot-collector-nginx
@@ -1,68 +0,0 @@
apiVersion: opentelemetry.io/v1alpha1
kind: OpenTelemetryCollector
metadata:
name: adot
spec:
image: public.ecr.aws/aws-observability/aws-otel-collector:v0.21.1
mode: deployment
serviceAccount: adot-collector-nginx
config: |
receivers:
prometheus:
config:
global:
scrape_interval: {{ .Values.scrapeInterval }}
scrape_timeout: {{ .Values.scrapeTimeout }}
scrape_configs:
- job_name: 'kubernetes-pod-nginx'
sample_limit: {{ .Values.scrapeSampleLimit }}
metrics_path: /{{ .Values.prometheusMetricsEndpoint }}
kubernetes_sd_configs:
- role: pod
relabel_configs:
- source_labels: [ __address__ ]
action: keep
regex: '.*:10254$'
- source_labels: [__meta_kubernetes_pod_container_name]
target_label: container
action: replace
- source_labels: [__meta_kubernetes_pod_node_name]
target_label: host
action: replace
- source_labels: [__meta_kubernetes_namespace]
target_label: namespace
action: replace
metric_relabel_configs:
- source_labels: [__name__]
regex: 'go_memstats.*'
action: drop
- source_labels: [__name__]
regex: 'go_gc.*'
action: drop
- source_labels: [__name__]
regex: 'go_threads'
action: drop
- regex: exported_host
action: labeldrop
exporters:
prometheusremotewrite:
endpoint: {{ .Values.ampurl }}
auth:
authenticator: sigv4auth
logging:
loglevel: info
extensions:
sigv4auth:
region: {{ .Values.region }}
service: "aps"
health_check:
pprof:
endpoint: :1888
zpages:
endpoint: :55679
service:
extensions: [pprof, zpages, health_check, sigv4auth]
pipelines:
metrics:
receivers: [prometheus]
exporters: [logging, prometheusremotewrite]
@@ -1,7 +0,0 @@
ampurl: ${amp_url}
region: ${region}
prometheusMetricsEndpoint: ${prometheus_metrics_endpoint}
prometheusMetricsPort: ${prometheus_metrics_port}
scrapeInterval: ${scrape_interval}
scrapeTimeout: ${scrape_timeout}
scrapeSampleLimit: ${scrape_sample_limit}
-69
View File
@@ -1,69 +0,0 @@
variable "eks_cluster_id" {
description = "EKS Cluster Id"
type = string
}
variable "helm_config" {
description = "Helm Config for Prometheus"
type = any
default = {}
}
variable "irsa_iam_role_path" {
description = "IAM role path for IRSA roles"
type = string
default = "/"
}
variable "irsa_iam_permissions_boundary" {
description = "IAM permissions boundary for IRSA roles"
type = string
default = null
}
variable "managed_prometheus_workspace_endpoint" {
description = "Amazon Managed Prometheus Workspace Endpoint"
type = string
default = ""
}
variable "managed_prometheus_workspace_id" {
description = "Amazon Managed Prometheus Workspace ID"
type = string
default = null
}
variable "managed_prometheus_workspace_region" {
description = "Amazon Managed Prometheus Workspace's Region"
type = string
default = null
}
variable "dashboards_folder_id" {
type = string
description = "Grafana folder ID for automatic dashboards"
}
variable "enable_alerting_rules" {
type = bool
default = true
description = "Enables or disables Managed Prometheus alerting rules"
}
variable "enable_dashboards" {
type = bool
description = "Enables or disables curated dashboards"
default = true
}
variable "config" {
description = "Helm Config for Prometheus"
type = any
default = {}
}
variable "tags" {
description = "Additional tags (e.g. `map('BusinessUnit`,`XYZ`)"
type = map(string)
default = {}
}
+5 -10
View File
@@ -1,18 +1,8 @@
output "eks_cluster_id" {
description = "EKS Cluster Id"
value = var.eks_cluster_id
}
output "aws_region" { output "aws_region" {
description = "AWS Region" description = "AWS Region"
value = var.aws_region value = var.aws_region
} }
output "eks_cluster_version" {
description = "EKS Cluster version"
value = data.aws_eks_cluster.eks_cluster.version
}
output "managed_prometheus_workspace_endpoint" { output "managed_prometheus_workspace_endpoint" {
description = "Amazon Managed Prometheus workspace endpoint" description = "Amazon Managed Prometheus workspace endpoint"
value = local.amp_ws_endpoint value = local.amp_ws_endpoint
@@ -42,3 +32,8 @@ output "grafana_dashboards_folder_id" {
description = "Grafana folder ID for automatic dashboards. Required by workload modules" description = "Grafana folder ID for automatic dashboards. Required by workload modules"
value = grafana_folder.this.id value = grafana_folder.this.id
} }
output "grafana_prometheus_datasource_test" {
description = "Grafana save & test URL for Amazon Managed Prometheus workspace"
value = "${local.amg_ws_endpoint}/datasources/edit/${grafana_data_source.amp.uid}"
}
-29
View File
@@ -1,37 +1,8 @@
variable "eks_cluster_id" {
description = "Name of the EKS cluster"
type = string
}
variable "aws_region" { variable "aws_region" {
description = "AWS Region" description = "AWS Region"
type = string type = string
} }
variable "irsa_iam_role_path" {
description = "IAM role path for IRSA roles"
type = string
default = "/"
}
variable "irsa_iam_permissions_boundary" {
description = "IAM permissions boundary for IRSA roles"
type = string
default = null
}
variable "enable_amazon_eks_adot" {
description = "Enables the ADOT Operator on the EKS Cluster"
type = bool
default = true
}
variable "enable_cert_manager" {
description = "Allow reusing an existing installation of cert-manager"
type = bool
default = true
}
variable "enable_managed_prometheus" { variable "enable_managed_prometheus" {
description = "Creates a new Amazon Managed Service for Prometheus Workspace" description = "Creates a new Amazon Managed Service for Prometheus Workspace"
type = bool type = bool