mirror of
https://github.com/storytold/terraform-aws-observability-accelerator.git
synced 2026-10-09 00:09:43 +00:00
Java dev (#52)
* Update Java module * Add java example * Updating Readme * Updating Readme Images * Pre-commit * Fixing readme screenshot * Update README.md Co-authored-by: Rodrigue Koffi <bonclay7@users.noreply.github.com>
This commit is contained in:
@@ -0,0 +1,240 @@
|
|||||||
|
# Existing Cluster with the AWS Observability accelerator base module and Java monitoring
|
||||||
|
|
||||||
|
|
||||||
|
This example demonstrates how to use the AWS Observability Accelerator Terraform
|
||||||
|
modules with Java monitoring enabled.
|
||||||
|
The current example deploys the [AWS Distro for OpenTelemetry Operator](https://docs.aws.amazon.com/eks/latest/userguide/opentelemetry.html) for Amazon EKS with its requirements and make use of existing
|
||||||
|
Amazon Managed Service for Prometheus and Amazon Managed Grafana workspaces.
|
||||||
|
|
||||||
|
It is based on the `java module`, one of our [workloads modules](../../modules/workloads/)
|
||||||
|
to provide an existing EKS cluster with an OpenTelemetry collector,
|
||||||
|
curated Grafana dashboards, Prometheus alerting and recording rules with multiple
|
||||||
|
configuration options on the cluster infrastructure.
|
||||||
|
|
||||||
|
|
||||||
|
## Prerequisites
|
||||||
|
|
||||||
|
Ensure that you have the following tools installed locally:
|
||||||
|
|
||||||
|
1. [aws cli v2](https://docs.aws.amazon.com/cli/latest/userguide/getting-started-install.html)
|
||||||
|
2. [kubectl](https://kubernetes.io/docs/tasks/tools/)
|
||||||
|
3. [terraform](https://learn.hashicorp.com/tutorials/terraform/install-cli)
|
||||||
|
|
||||||
|
|
||||||
|
## Setup
|
||||||
|
|
||||||
|
This example uses a local terraform state. If you need states to be saved remotely,
|
||||||
|
on Amazon S3 for example, visit the [terraform remote states](https://www.terraform.io/language/state/remote) documentation
|
||||||
|
|
||||||
|
1. Clone the repo using the command below
|
||||||
|
|
||||||
|
```
|
||||||
|
git clone https://github.com/aws-observability/terraform-aws-observability-accelerator.git
|
||||||
|
```
|
||||||
|
|
||||||
|
2. Initialize terraform
|
||||||
|
|
||||||
|
```console
|
||||||
|
cd examples/existing-cluster-java
|
||||||
|
terraform init
|
||||||
|
```
|
||||||
|
|
||||||
|
3. AWS Region
|
||||||
|
|
||||||
|
Specify the AWS Region where the resources will be deployed. Edit the `terraform.tfvars` file and modify `aws_region="..."`. You can also use environement variables `export TF_VAR_aws_region=xxx`.
|
||||||
|
|
||||||
|
4. Amazon EKS Cluster
|
||||||
|
|
||||||
|
To run this example, you need to provide your EKS cluster name.
|
||||||
|
If you don't have a cluster ready, visit [this example](../eks-cluster-with-vpc)
|
||||||
|
first to create a new one.
|
||||||
|
|
||||||
|
Add your cluster name for `eks_cluster_id="..."` to the `terraform.tfvars` or use an environment variable `export TF_VAR_eks_cluster_id=xxx`.
|
||||||
|
|
||||||
|
5. Amazon Managed Service for Prometheus workspace (optional)
|
||||||
|
|
||||||
|
If you have an existing workspace, add `managed_prometheus_workspace_id=ws-xxx`
|
||||||
|
or use an environment variable `export TF_VAR_managed_prometheus_workspace_id=ws-xxx`.
|
||||||
|
|
||||||
|
If you don't specify anything a new workspace will be created for you.
|
||||||
|
|
||||||
|
6. Amazon Managed Grafana workspace
|
||||||
|
|
||||||
|
If you have an existing workspace, create an environment variable `export TF_VAR_managed_grafana_workspace_id=g-xxx`.
|
||||||
|
|
||||||
|
7. <a name="apikey"></a> Grafana API Key
|
||||||
|
|
||||||
|
Amazon Managed Service for Grafana provides a control plane API for generating Grafana API keys. We will provide to Terraform
|
||||||
|
a short lived API key to run the `apply` or `destroy` command.
|
||||||
|
Ensure you have necessary IAM permissions (`CreateWorkspaceApiKey, DeleteWorkspaceApiKey`)
|
||||||
|
|
||||||
|
```sh
|
||||||
|
export TF_VAR_grafana_api_key=`aws grafana create-workspace-api-key --key-name "observability-accelerator-$(date +%s)" --key-role ADMIN --seconds-to-live 1200 --workspace-id $TF_VAR_managed_grafana_workspace_id --query key --output text`
|
||||||
|
```
|
||||||
|
|
||||||
|
## Deploy
|
||||||
|
|
||||||
|
```sh
|
||||||
|
terraform apply -var-file=terraform.tfvars
|
||||||
|
```
|
||||||
|
|
||||||
|
or if you had only setup environment variables, run
|
||||||
|
|
||||||
|
```sh
|
||||||
|
terraform apply
|
||||||
|
```
|
||||||
|
|
||||||
|
## Visualization
|
||||||
|
|
||||||
|
1. Prometheus datasource on Grafana
|
||||||
|
|
||||||
|
Open your Grafana workspace and under Configuration -> Data sources, you will see `aws-observability-accelerator`. Open and click `Save & test`. You will then see a notification confirming that the Amazon Managed Service for Prometheus workspace is ready to be used on Grafana.
|
||||||
|
|
||||||
|
2. Grafana dashboards
|
||||||
|
|
||||||
|
Go to the Dashboards panel of your Grafana workspace. There will be a folder called `Observability Accelerator Dashboards`
|
||||||
|
|
||||||
|
<img width="832" alt="image" src="https://user-images.githubusercontent.com/97046295/194903648-57c55d30-6f90-4b03-9eb6-577aaba7dc22.png">
|
||||||
|
|
||||||
|
Open the "Java/JMX" dashboard to view its visualization
|
||||||
|
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
|
||||||
|
2. Amazon Managed Service for Prometheus rules and alerts
|
||||||
|
|
||||||
|
Open the Amazon Managed Service for Prometheus console and view the details of your workspace. Under the `Rules management` tab, you will find new rules deployed.
|
||||||
|
|
||||||
|
<img width="1314" alt="image" src="https://user-images.githubusercontent.com/97046295/194904104-09a28577-d149-478e-b0a1-dc21cb7effc1.png">
|
||||||
|
|
||||||
|
|
||||||
|
To setup your alert receiver, with Amazon SNS, follow [this documentation](https://docs.aws.amazon.com/prometheus/latest/userguide/AMP-alertmanager-receiver.html)
|
||||||
|
|
||||||
|
|
||||||
|
## Deploy an Example Java Application
|
||||||
|
|
||||||
|
In this section we will reuse an example from the AWS OpenTelemetry collector [repository](https://github.com/aws-observability/aws-otel-collector/blob/main/docs/developers/container-insights-eks-jmx.md). For convenience, the steps can be found below.
|
||||||
|
|
||||||
|
1. Clone [this repository](https://github.com/aws-observability/aws-otel-test-framework) and navigate to the `sample-apps/jmx/` directory.
|
||||||
|
|
||||||
|
2. Authenticate to Amazon ECR
|
||||||
|
|
||||||
|
```sh
|
||||||
|
export AWS_ACCOUNT_ID=`aws sts get-caller-identity --query Account --output text`
|
||||||
|
export AWS_REGION={region}
|
||||||
|
aws ecr get-login-password --region $AWS_REGION | docker login --username AWS --password-stdin $AWS_ACCOUNT_ID.dkr.ecr.$AWS_REGION.amazonaws.com
|
||||||
|
```
|
||||||
|
|
||||||
|
3. Create an Amazon ECR repository
|
||||||
|
|
||||||
|
```sh
|
||||||
|
aws ecr create-repository --repository-name prometheus-sample-tomcat-jmx \
|
||||||
|
--image-scanning-configuration scanOnPush=true \
|
||||||
|
--region $AWS_REGION
|
||||||
|
```
|
||||||
|
|
||||||
|
4. Build Docker image and push to ECR.
|
||||||
|
|
||||||
|
```sh
|
||||||
|
docker build -t $AWS_ACCOUNT_ID.dkr.ecr.$AWS_REGION.amazonaws.com/prometheus-sample-tomcat-jmx:latest .
|
||||||
|
docker push $AWS_ACCOUNT_ID.dkr.ecr.$AWS_REGION.amazonaws.com/prometheus-sample-tomcat-jmx:latest
|
||||||
|
```
|
||||||
|
|
||||||
|
5. Install sample application
|
||||||
|
|
||||||
|
```sh
|
||||||
|
export SAMPLE_TRAFFIC_NAMESPACE=javajmx-sample
|
||||||
|
curl https://raw.githubusercontent.com/aws-observability/aws-otel-test-framework/terraform/sample-apps/jmx/examples/prometheus-metrics-sample.yaml > metrics-sample.yaml
|
||||||
|
sed -i "s/{{aws_account_id}}/$AWS_ACCOUNT_ID/g" metrics-sample.yaml
|
||||||
|
sed -i "s/{{region}}/$AWS_REGION/g" metrics-sample.yaml
|
||||||
|
sed -i "s/{{namespace}}/$SAMPLE_TRAFFIC_NAMESPACE/g" metrics-sample.yaml
|
||||||
|
kubectl apply -f metrics-sample.yaml
|
||||||
|
```
|
||||||
|
|
||||||
|
Verify that the sample application is running:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
kubectl get pods -n $SAMPLE_TRAFFIC_NAMESPACE
|
||||||
|
|
||||||
|
NAME READY STATUS RESTARTS AGE
|
||||||
|
tomcat-bad-traffic-generator 1/1 Running 0 11s
|
||||||
|
tomcat-example-7958666589-2q755 0/1 ContainerCreating 0 11s
|
||||||
|
tomcat-traffic-generator 1/1 Running 0 11s
|
||||||
|
```
|
||||||
|
|
||||||
|
## Advanced configuration
|
||||||
|
|
||||||
|
1. Cross-region Amazon Managed Prometheus workspace
|
||||||
|
|
||||||
|
If your existing Amazon Managed Prometheus workspace is in another AWS Region,
|
||||||
|
add this `managed_prometheus_region=xxx` and `managed_prometheus_workspace_id=ws-xxx`.
|
||||||
|
|
||||||
|
2. Cross-region Amazon Managed Grafana workspace
|
||||||
|
|
||||||
|
If your existing Amazon Managed Prometheus workspace is in another AWS Region,
|
||||||
|
add this `managed_prometheus_region=xxx` and `managed_prometheus_workspace_id=ws-xxx`.
|
||||||
|
|
||||||
|
## Destroy resources
|
||||||
|
|
||||||
|
If you leave this stack running, you will continue to incur charges. To remove all resources
|
||||||
|
created by Terraform, [refresh your Grafana API key](#apikey) and run:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
terraform destroy -var-file=terraform.tfvars
|
||||||
|
```
|
||||||
|
|
||||||
|
|
||||||
|
<!-- BEGINNING OF PRE-COMMIT-TERRAFORM DOCS HOOK -->
|
||||||
|
## Requirements
|
||||||
|
|
||||||
|
| Name | Version |
|
||||||
|
|------|---------|
|
||||||
|
| <a name="requirement_terraform"></a> [terraform](#requirement\_terraform) | >= 1.1.0, < 1.3.0 |
|
||||||
|
| <a name="requirement_aws"></a> [aws](#requirement\_aws) | >= 4.0.0 |
|
||||||
|
| <a name="requirement_grafana"></a> [grafana](#requirement\_grafana) | >= 1.25.0 |
|
||||||
|
| <a name="requirement_helm"></a> [helm](#requirement\_helm) | >= 2.4.1 |
|
||||||
|
| <a name="requirement_kubectl"></a> [kubectl](#requirement\_kubectl) | >= 1.14 |
|
||||||
|
| <a name="requirement_kubernetes"></a> [kubernetes](#requirement\_kubernetes) | >= 2.10 |
|
||||||
|
|
||||||
|
## Providers
|
||||||
|
|
||||||
|
| Name | Version |
|
||||||
|
|------|---------|
|
||||||
|
| <a name="provider_aws"></a> [aws](#provider\_aws) | >= 4.0.0 |
|
||||||
|
|
||||||
|
## Modules
|
||||||
|
|
||||||
|
| Name | Source | Version |
|
||||||
|
|------|--------|---------|
|
||||||
|
| <a name="module_eks_observability_accelerator"></a> [eks\_observability\_accelerator](#module\_eks\_observability\_accelerator) | ../../ | n/a |
|
||||||
|
| <a name="module_workloads_java"></a> [workloads\_java](#module\_workloads\_java) | ../../modules/workloads/java | n/a |
|
||||||
|
|
||||||
|
## Resources
|
||||||
|
|
||||||
|
| Name | Type |
|
||||||
|
|------|------|
|
||||||
|
| [aws_eks_cluster.this](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/data-sources/eks_cluster) | data source |
|
||||||
|
| [aws_eks_cluster_auth.this](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/data-sources/eks_cluster_auth) | data source |
|
||||||
|
|
||||||
|
## Inputs
|
||||||
|
|
||||||
|
| Name | Description | Type | Default | Required |
|
||||||
|
|------|-------------|------|---------|:--------:|
|
||||||
|
| <a name="input_aws_region"></a> [aws\_region](#input\_aws\_region) | AWS Region | `string` | n/a | yes |
|
||||||
|
| <a name="input_eks_cluster_id"></a> [eks\_cluster\_id](#input\_eks\_cluster\_id) | Name of the EKS cluster | `string` | n/a | yes |
|
||||||
|
| <a name="input_grafana_api_key"></a> [grafana\_api\_key](#input\_grafana\_api\_key) | API key for authorizing the Grafana provider to make changes to Amazon Managed Grafana | `string` | `""` | no |
|
||||||
|
| <a name="input_managed_grafana_workspace_id"></a> [managed\_grafana\_workspace\_id](#input\_managed\_grafana\_workspace\_id) | Amazon Managed Grafana Workspace ID | `string` | `""` | no |
|
||||||
|
| <a name="input_managed_prometheus_workspace_id"></a> [managed\_prometheus\_workspace\_id](#input\_managed\_prometheus\_workspace\_id) | Amazon Managed Service for Prometheus Workspace ID | `string` | `""` | no |
|
||||||
|
|
||||||
|
## Outputs
|
||||||
|
|
||||||
|
| Name | Description |
|
||||||
|
|------|-------------|
|
||||||
|
| <a name="output_aws_region"></a> [aws\_region](#output\_aws\_region) | AWS Region |
|
||||||
|
| <a name="output_eks_cluster_id"></a> [eks\_cluster\_id](#output\_eks\_cluster\_id) | EKS Cluster Id |
|
||||||
|
| <a name="output_eks_cluster_version"></a> [eks\_cluster\_version](#output\_eks\_cluster\_version) | EKS Cluster version |
|
||||||
|
| <a name="output_grafana_dashboard_urls"></a> [grafana\_dashboard\_urls](#output\_grafana\_dashboard\_urls) | URLs for dashboards created |
|
||||||
|
| <a name="output_managed_prometheus_workspace_endpoint"></a> [managed\_prometheus\_workspace\_endpoint](#output\_managed\_prometheus\_workspace\_endpoint) | Amazon Managed Prometheus workspace endpoint |
|
||||||
|
| <a name="output_managed_prometheus_workspace_id"></a> [managed\_prometheus\_workspace\_id](#output\_managed\_prometheus\_workspace\_id) | Amazon Managed Prometheus workspace ID |
|
||||||
|
<!-- END OF PRE-COMMIT-TERRAFORM DOCS HOOK -->
|
||||||
@@ -0,0 +1,100 @@
|
|||||||
|
provider "aws" {
|
||||||
|
region = local.region
|
||||||
|
}
|
||||||
|
|
||||||
|
data "aws_eks_cluster_auth" "this" {
|
||||||
|
name = var.eks_cluster_id
|
||||||
|
}
|
||||||
|
|
||||||
|
data "aws_eks_cluster" "this" {
|
||||||
|
name = var.eks_cluster_id
|
||||||
|
}
|
||||||
|
|
||||||
|
provider "kubernetes" {
|
||||||
|
host = local.eks_cluster_endpoint
|
||||||
|
cluster_ca_certificate = base64decode(data.aws_eks_cluster.this.certificate_authority[0].data)
|
||||||
|
token = data.aws_eks_cluster_auth.this.token
|
||||||
|
}
|
||||||
|
|
||||||
|
provider "helm" {
|
||||||
|
kubernetes {
|
||||||
|
host = local.eks_cluster_endpoint
|
||||||
|
cluster_ca_certificate = base64decode(data.aws_eks_cluster.this.certificate_authority[0].data)
|
||||||
|
token = data.aws_eks_cluster_auth.this.token
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
locals {
|
||||||
|
region = var.aws_region
|
||||||
|
eks_cluster_endpoint = data.aws_eks_cluster.this.endpoint
|
||||||
|
create_new_workspace = var.managed_prometheus_workspace_id == "" ? true : false
|
||||||
|
tags = {
|
||||||
|
Source = "github.com/aws-observability/terraform-aws-observability-accelerator"
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
# deploys the base module
|
||||||
|
module "eks_observability_accelerator" {
|
||||||
|
# source = "aws-observability/terrarom-aws-observability-accelerator"
|
||||||
|
source = "../../"
|
||||||
|
|
||||||
|
aws_region = var.aws_region
|
||||||
|
eks_cluster_id = var.eks_cluster_id
|
||||||
|
|
||||||
|
# deploys AWS Distro for OpenTelemetry operator into the cluster
|
||||||
|
enable_amazon_eks_adot = true
|
||||||
|
|
||||||
|
# reusing existing certificate manager? defaults to true
|
||||||
|
enable_cert_manager = true
|
||||||
|
|
||||||
|
# creates a new Amazon Managed Prometheus workspace, defaults to true
|
||||||
|
enable_managed_prometheus = local.create_new_workspace
|
||||||
|
|
||||||
|
# reusing existing Amazon Managed Prometheus if specified
|
||||||
|
managed_prometheus_workspace_id = var.managed_prometheus_workspace_id
|
||||||
|
managed_prometheus_workspace_region = null # defaults to the current region, useful for cross region scenarios (same account)
|
||||||
|
|
||||||
|
# sets up the Amazon Managed Prometheus alert manager at the workspace level
|
||||||
|
enable_alertmanager = true
|
||||||
|
|
||||||
|
# reusing existing Amazon Managed Grafana workspace
|
||||||
|
enable_managed_grafana = false
|
||||||
|
managed_grafana_workspace_id = var.managed_grafana_workspace_id
|
||||||
|
grafana_api_key = var.grafana_api_key
|
||||||
|
|
||||||
|
tags = local.tags
|
||||||
|
}
|
||||||
|
|
||||||
|
# https://www.terraform.io/language/modules/develop/providers
|
||||||
|
# A module intended to be called by one or more other modules must not contain
|
||||||
|
# any provider blocks.
|
||||||
|
# This allows forcing dependency between base and workloads module
|
||||||
|
provider "grafana" {
|
||||||
|
url = module.eks_observability_accelerator.managed_grafana_workspace_endpoint
|
||||||
|
auth = var.grafana_api_key
|
||||||
|
}
|
||||||
|
|
||||||
|
module "workloads_java" {
|
||||||
|
source = "../../modules/workloads/java"
|
||||||
|
|
||||||
|
eks_cluster_id = module.eks_observability_accelerator.eks_cluster_id
|
||||||
|
|
||||||
|
dashboards_folder_id = module.eks_observability_accelerator.grafana_dashboards_folder_id
|
||||||
|
managed_prometheus_workspace_id = module.eks_observability_accelerator.managed_prometheus_workspace_id
|
||||||
|
|
||||||
|
managed_prometheus_workspace_endpoint = module.eks_observability_accelerator.managed_prometheus_workspace_endpoint
|
||||||
|
managed_prometheus_workspace_region = module.eks_observability_accelerator.managed_prometheus_workspace_region
|
||||||
|
|
||||||
|
# optional, defaults to 60s interval and 15s timeout
|
||||||
|
prometheus_config = {
|
||||||
|
global_scrape_interval = "60s"
|
||||||
|
global_scrape_timeout = "15s"
|
||||||
|
scrape_sample_limit = 2000
|
||||||
|
}
|
||||||
|
|
||||||
|
tags = local.tags
|
||||||
|
|
||||||
|
depends_on = [
|
||||||
|
module.eks_observability_accelerator
|
||||||
|
]
|
||||||
|
}
|
||||||
@@ -0,0 +1,29 @@
|
|||||||
|
output "eks_cluster_id" {
|
||||||
|
description = "EKS Cluster Id"
|
||||||
|
value = module.eks_observability_accelerator.eks_cluster_id
|
||||||
|
}
|
||||||
|
|
||||||
|
output "aws_region" {
|
||||||
|
description = "AWS Region"
|
||||||
|
value = module.eks_observability_accelerator.aws_region
|
||||||
|
}
|
||||||
|
|
||||||
|
output "eks_cluster_version" {
|
||||||
|
description = "EKS Cluster version"
|
||||||
|
value = module.eks_observability_accelerator.eks_cluster_version
|
||||||
|
}
|
||||||
|
|
||||||
|
output "managed_prometheus_workspace_endpoint" {
|
||||||
|
description = "Amazon Managed Prometheus workspace endpoint"
|
||||||
|
value = module.eks_observability_accelerator.managed_prometheus_workspace_endpoint
|
||||||
|
}
|
||||||
|
|
||||||
|
output "managed_prometheus_workspace_id" {
|
||||||
|
description = "Amazon Managed Prometheus workspace ID"
|
||||||
|
value = module.eks_observability_accelerator.managed_prometheus_workspace_id
|
||||||
|
}
|
||||||
|
|
||||||
|
output "grafana_dashboard_urls" {
|
||||||
|
description = "URLs for dashboards created"
|
||||||
|
value = module.workloads_java.grafana_dashboard_urls
|
||||||
|
}
|
||||||
@@ -0,0 +1,24 @@
|
|||||||
|
variable "eks_cluster_id" {
|
||||||
|
description = "Name of the EKS cluster"
|
||||||
|
type = string
|
||||||
|
}
|
||||||
|
variable "aws_region" {
|
||||||
|
description = "AWS Region"
|
||||||
|
type = string
|
||||||
|
}
|
||||||
|
variable "managed_prometheus_workspace_id" {
|
||||||
|
description = "Amazon Managed Service for Prometheus Workspace ID"
|
||||||
|
type = string
|
||||||
|
default = ""
|
||||||
|
}
|
||||||
|
variable "managed_grafana_workspace_id" {
|
||||||
|
description = "Amazon Managed Grafana Workspace ID"
|
||||||
|
type = string
|
||||||
|
default = ""
|
||||||
|
}
|
||||||
|
variable "grafana_api_key" {
|
||||||
|
description = "API key for authorizing the Grafana provider to make changes to Amazon Managed Grafana"
|
||||||
|
type = string
|
||||||
|
default = ""
|
||||||
|
sensitive = true
|
||||||
|
}
|
||||||
@@ -0,0 +1,34 @@
|
|||||||
|
terraform {
|
||||||
|
required_version = ">= 1.1.0, < 1.3.0"
|
||||||
|
|
||||||
|
required_providers {
|
||||||
|
aws = {
|
||||||
|
source = "hashicorp/aws"
|
||||||
|
version = ">= 4.0.0"
|
||||||
|
}
|
||||||
|
kubernetes = {
|
||||||
|
source = "hashicorp/kubernetes"
|
||||||
|
version = ">= 2.10"
|
||||||
|
}
|
||||||
|
kubectl = {
|
||||||
|
source = "gavinbunney/kubectl"
|
||||||
|
version = ">= 1.14"
|
||||||
|
}
|
||||||
|
helm = {
|
||||||
|
source = "hashicorp/helm"
|
||||||
|
version = ">= 2.4.1"
|
||||||
|
}
|
||||||
|
grafana = {
|
||||||
|
source = "grafana/grafana"
|
||||||
|
version = ">= 1.25.0"
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
# ## Used for end-to-end testing on project; update to suit your needs
|
||||||
|
# backend "s3" {
|
||||||
|
# bucket = "observability-accelerator-terraform-states"
|
||||||
|
# region = "us-west-2"
|
||||||
|
# key = "e2e/existing-cluster-with-base-and-infra/terraform.tfstate"
|
||||||
|
# }
|
||||||
|
|
||||||
|
}
|
||||||
@@ -29,29 +29,40 @@ This module provides monitoring for Java based workloads with the following reso
|
|||||||
|
|
||||||
| Name | Source | Version |
|
| Name | Source | Version |
|
||||||
|------|--------|---------|
|
|------|--------|---------|
|
||||||
| <a name="module_helm_addon"></a> [helm\_addon](#module\_helm\_addon) | github.com/aws-ia/terraform-aws-eks-blueprints//modules/kubernetes-addons/helm-addon | v4.8.1 |
|
| <a name="module_helm_addon"></a> [helm\_addon](#module\_helm\_addon) | github.com/aws-ia/terraform-aws-eks-blueprints//modules/kubernetes-addons/helm-addon | v4.12.1 |
|
||||||
|
|
||||||
## Resources
|
## Resources
|
||||||
|
|
||||||
| Name | Type |
|
| Name | Type |
|
||||||
|------|------|
|
|------|------|
|
||||||
| [aws_prometheus_rule_group_namespace.this](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/resources/prometheus_rule_group_namespace) | resource |
|
| [aws_prometheus_rule_group_namespace.alerting_rules](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/resources/prometheus_rule_group_namespace) | resource |
|
||||||
|
| [aws_prometheus_rule_group_namespace.recording_rules](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/resources/prometheus_rule_group_namespace) | resource |
|
||||||
| [grafana_dashboard.this](https://registry.terraform.io/providers/grafana/grafana/latest/docs/resources/dashboard) | resource |
|
| [grafana_dashboard.this](https://registry.terraform.io/providers/grafana/grafana/latest/docs/resources/dashboard) | resource |
|
||||||
|
| [aws_caller_identity.current](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/data-sources/caller_identity) | data source |
|
||||||
|
| [aws_eks_cluster.eks_cluster](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/data-sources/eks_cluster) | data source |
|
||||||
| [aws_partition.current](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/data-sources/partition) | data source |
|
| [aws_partition.current](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/data-sources/partition) | data source |
|
||||||
|
| [aws_region.current](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/data-sources/region) | data source |
|
||||||
|
|
||||||
## Inputs
|
## Inputs
|
||||||
|
|
||||||
| Name | Description | Type | Default | Required |
|
| Name | Description | Type | Default | Required |
|
||||||
|------|-------------|------|---------|:--------:|
|
|------|-------------|------|---------|:--------:|
|
||||||
| <a name="input_addon_context"></a> [addon\_context](#input\_addon\_context) | Input configuration for the addon | <pre>object({<br> aws_caller_identity_account_id = string<br> aws_caller_identity_arn = string<br> aws_eks_cluster_endpoint = string<br> aws_partition_id = string<br> aws_region_name = string<br> eks_cluster_id = string<br> eks_oidc_issuer_url = string<br> eks_oidc_provider_arn = string<br> irsa_iam_permissions_boundary = string<br> irsa_iam_role_path = string<br> tags = map(string)<br> })</pre> | n/a | yes |
|
|
||||||
| <a name="input_amp_endpoint"></a> [amp\_endpoint](#input\_amp\_endpoint) | Amazon Managed Prometheus endpoint | `string` | n/a | yes |
|
|
||||||
| <a name="input_amp_id"></a> [amp\_id](#input\_amp\_id) | Managed Prometheus workspace id | `string` | n/a | yes |
|
|
||||||
| <a name="input_amp_region"></a> [amp\_region](#input\_amp\_region) | Amazon Managed Prometheus Workspace's Region | `string` | `null` | no |
|
|
||||||
| <a name="input_dashboards_folder_id"></a> [dashboards\_folder\_id](#input\_dashboards\_folder\_id) | Grafana folder ID for automatic dashboards | `string` | n/a | yes |
|
| <a name="input_dashboards_folder_id"></a> [dashboards\_folder\_id](#input\_dashboards\_folder\_id) | Grafana folder ID for automatic dashboards | `string` | n/a | yes |
|
||||||
| <a name="input_enable_recording_rules"></a> [enable\_recording\_rules](#input\_enable\_recording\_rules) | Enable AMP recording rules | `bool` | `true` | no |
|
| <a name="input_eks_cluster_id"></a> [eks\_cluster\_id](#input\_eks\_cluster\_id) | EKS Cluster Id | `string` | n/a | yes |
|
||||||
|
| <a name="input_enable_alerting_rules"></a> [enable\_alerting\_rules](#input\_enable\_alerting\_rules) | Enables or disables Managed Prometheus alerting rules | `bool` | `true` | no |
|
||||||
|
| <a name="input_enable_recording_rules"></a> [enable\_recording\_rules](#input\_enable\_recording\_rules) | Enables or disables Managed Prometheus recording rules. Disabling this might affect some data in the dashboards | `bool` | `true` | no |
|
||||||
| <a name="input_helm_config"></a> [helm\_config](#input\_helm\_config) | Helm Config for Prometheus | `any` | `{}` | no |
|
| <a name="input_helm_config"></a> [helm\_config](#input\_helm\_config) | Helm Config for Prometheus | `any` | `{}` | no |
|
||||||
|
| <a name="input_irsa_iam_permissions_boundary"></a> [irsa\_iam\_permissions\_boundary](#input\_irsa\_iam\_permissions\_boundary) | IAM permissions boundary for IRSA roles | `string` | `""` | no |
|
||||||
|
| <a name="input_irsa_iam_role_path"></a> [irsa\_iam\_role\_path](#input\_irsa\_iam\_role\_path) | IAM role path for IRSA roles | `string` | `"/"` | no |
|
||||||
|
| <a name="input_managed_prometheus_workspace_endpoint"></a> [managed\_prometheus\_workspace\_endpoint](#input\_managed\_prometheus\_workspace\_endpoint) | Amazon Managed Prometheus Workspace Endpoint | `string` | `null` | no |
|
||||||
|
| <a name="input_managed_prometheus_workspace_id"></a> [managed\_prometheus\_workspace\_id](#input\_managed\_prometheus\_workspace\_id) | Amazon Managed Prometheus Workspace ID | `string` | `null` | no |
|
||||||
|
| <a name="input_managed_prometheus_workspace_region"></a> [managed\_prometheus\_workspace\_region](#input\_managed\_prometheus\_workspace\_region) | Amazon Managed Prometheus Workspace's Region | `string` | `null` | no |
|
||||||
|
| <a name="input_prometheus_config"></a> [prometheus\_config](#input\_prometheus\_config) | Controls default values such as scrape interval, timeouts and ports globally | <pre>object({<br> global_scrape_interval = string<br> global_scrape_timeout = string<br> scrape_sample_limit = number<br> })</pre> | <pre>{<br> "global_scrape_interval": "60s",<br> "global_scrape_timeout": "15s",<br> "scrape_sample_limit": 1000<br>}</pre> | no |
|
||||||
|
| <a name="input_tags"></a> [tags](#input\_tags) | Additional tags (e.g. `map('BusinessUnit`,`XYZ`) | `map(string)` | `{}` | no |
|
||||||
|
|
||||||
## Outputs
|
## Outputs
|
||||||
|
|
||||||
No outputs.
|
| Name | Description |
|
||||||
|
|------|-------------|
|
||||||
|
| <a name="output_grafana_dashboard_urls"></a> [grafana\_dashboard\_urls](#output\_grafana\_dashboard\_urls) | URLs for dashboards created |
|
||||||
<!-- END OF PRE-COMMIT-TERRAFORM DOCS HOOK -->
|
<!-- END OF PRE-COMMIT-TERRAFORM DOCS HOOK -->
|
||||||
|
|||||||
@@ -0,0 +1,31 @@
|
|||||||
|
data "aws_partition" "current" {}
|
||||||
|
|
||||||
|
data "aws_caller_identity" "current" {}
|
||||||
|
|
||||||
|
data "aws_region" "current" {}
|
||||||
|
|
||||||
|
data "aws_eks_cluster" "eks_cluster" {
|
||||||
|
name = var.eks_cluster_id
|
||||||
|
}
|
||||||
|
|
||||||
|
locals {
|
||||||
|
name = "adot-collector-java"
|
||||||
|
namespace = try(var.helm_config.namespace, local.name)
|
||||||
|
|
||||||
|
eks_oidc_issuer_url = replace(data.aws_eks_cluster.eks_cluster.identity[0].oidc[0].issuer, "https://", "")
|
||||||
|
eks_cluster_endpoint = data.aws_eks_cluster.eks_cluster.endpoint
|
||||||
|
|
||||||
|
context = {
|
||||||
|
aws_caller_identity_account_id = data.aws_caller_identity.current.account_id
|
||||||
|
aws_caller_identity_arn = data.aws_caller_identity.current.arn
|
||||||
|
aws_eks_cluster_endpoint = local.eks_cluster_endpoint
|
||||||
|
aws_partition_id = data.aws_partition.current.partition
|
||||||
|
aws_region_name = data.aws_region.current.name
|
||||||
|
eks_cluster_id = var.eks_cluster_id
|
||||||
|
eks_oidc_issuer_url = local.eks_oidc_issuer_url
|
||||||
|
eks_oidc_provider_arn = "arn:${data.aws_partition.current.partition}:iam::${data.aws_caller_identity.current.account_id}:oidc-provider/${local.eks_oidc_issuer_url}"
|
||||||
|
tags = var.tags
|
||||||
|
irsa_iam_role_path = var.irsa_iam_role_path
|
||||||
|
irsa_iam_permissions_boundary = var.irsa_iam_permissions_boundary
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -1,13 +1,6 @@
|
|||||||
locals {
|
|
||||||
name = "adot-collector-java"
|
|
||||||
namespace = try(var.helm_config.namespace, local.name)
|
|
||||||
}
|
|
||||||
|
|
||||||
data "aws_partition" "current" {}
|
|
||||||
|
|
||||||
# deploys collector
|
# deploys collector
|
||||||
module "helm_addon" {
|
module "helm_addon" {
|
||||||
source = "github.com/aws-ia/terraform-aws-eks-blueprints//modules/kubernetes-addons/helm-addon?ref=v4.8.1"
|
source = "github.com/aws-ia/terraform-aws-eks-blueprints//modules/kubernetes-addons/helm-addon?ref=v4.12.1"
|
||||||
|
|
||||||
helm_config = merge(
|
helm_config = merge(
|
||||||
{
|
{
|
||||||
@@ -23,31 +16,31 @@ module "helm_addon" {
|
|||||||
set_values = [
|
set_values = [
|
||||||
{
|
{
|
||||||
name = "ampurl"
|
name = "ampurl"
|
||||||
value = "${var.amp_endpoint}api/v1/remote_write"
|
value = "${var.managed_prometheus_workspace_endpoint}api/v1/remote_write"
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
name = "region"
|
name = "region"
|
||||||
value = var.amp_region
|
value = var.managed_prometheus_workspace_region
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
name = "prometheusMetricsEndpoint"
|
name = "ekscluster"
|
||||||
value = "metrics"
|
value = local.context.eks_cluster_id
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
name = "prometheusMetricsPort"
|
name = "accountId"
|
||||||
value = 8888
|
value = local.context.aws_caller_identity_account_id
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
name = "scrapeInterval"
|
name = "globalScrapeInterval"
|
||||||
value = "15s"
|
value = var.prometheus_config.global_scrape_interval
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
name = "scrapeTimeout"
|
name = "globalScrapeTimeout"
|
||||||
value = "10s"
|
value = var.prometheus_config.global_scrape_timeout
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
name = "scrapeSampleLimit"
|
name = "scrapeSampleLimit"
|
||||||
value = 1000
|
value = var.prometheus_config.scrape_sample_limit
|
||||||
}
|
}
|
||||||
]
|
]
|
||||||
|
|
||||||
@@ -59,21 +52,30 @@ module "helm_addon" {
|
|||||||
irsa_iam_policies = ["arn:${data.aws_partition.current.partition}:iam::aws:policy/AmazonPrometheusRemoteWriteAccess"]
|
irsa_iam_policies = ["arn:${data.aws_partition.current.partition}:iam::aws:policy/AmazonPrometheusRemoteWriteAccess"]
|
||||||
}
|
}
|
||||||
|
|
||||||
addon_context = var.addon_context
|
addon_context = local.context
|
||||||
}
|
}
|
||||||
|
|
||||||
|
resource "aws_prometheus_rule_group_namespace" "recording_rules" {
|
||||||
resource "aws_prometheus_rule_group_namespace" "this" {
|
|
||||||
count = var.enable_recording_rules ? 1 : 0
|
count = var.enable_recording_rules ? 1 : 0
|
||||||
|
|
||||||
name = "java_rules"
|
name = "accelerator-java-rules"
|
||||||
workspace_id = var.amp_id
|
workspace_id = var.managed_prometheus_workspace_id
|
||||||
data = <<EOF
|
data = <<EOF
|
||||||
groups:
|
groups:
|
||||||
- name: default-metric
|
- name: default-metric
|
||||||
rules:
|
rules:
|
||||||
- record: metric:recording_rule
|
- record: metric:recording_rule
|
||||||
expr: avg(rate(container_cpu_usage_seconds_total[5m]))
|
expr: avg(rate(container_cpu_usage_seconds_total[5m]))
|
||||||
|
EOF
|
||||||
|
}
|
||||||
|
|
||||||
|
resource "aws_prometheus_rule_group_namespace" "alerting_rules" {
|
||||||
|
count = var.enable_alerting_rules ? 1 : 0
|
||||||
|
|
||||||
|
name = "accelerator-java-alerting"
|
||||||
|
workspace_id = var.managed_prometheus_workspace_id
|
||||||
|
data = <<EOF
|
||||||
|
groups:
|
||||||
- name: default-alert
|
- name: default-alert
|
||||||
rules:
|
rules:
|
||||||
- alert: metric:alerting_rule
|
- alert: metric:alerting_rule
|
||||||
|
|||||||
@@ -3,7 +3,7 @@ kind: OpenTelemetryCollector
|
|||||||
metadata:
|
metadata:
|
||||||
name: adot
|
name: adot
|
||||||
spec:
|
spec:
|
||||||
image: public.ecr.aws/aws-observability/aws-otel-collector:latest
|
image: public.ecr.aws/aws-observability/aws-otel-collector:v0.22.0
|
||||||
mode: deployment
|
mode: deployment
|
||||||
serviceAccount: adot-collector-java
|
serviceAccount: adot-collector-java
|
||||||
config: |
|
config: |
|
||||||
@@ -11,13 +11,15 @@ spec:
|
|||||||
prometheus:
|
prometheus:
|
||||||
config:
|
config:
|
||||||
global:
|
global:
|
||||||
scrape_interval: {{ .Values.scrapeInterval }}
|
scrape_interval: {{ .Values.globalScrapeInterval }}
|
||||||
scrape_timeout: {{ .Values.scrapeTimeout }}
|
scrape_timeout: {{ .Values.globalScrapeTimeout }}
|
||||||
|
external_labels:
|
||||||
|
cluster: {{ .Values.ekscluster }}
|
||||||
|
account_id: {{ .Values.accountId }}
|
||||||
|
region: {{ .Values.region }}
|
||||||
scrape_configs:
|
scrape_configs:
|
||||||
- job_name: 'kubernetes-pod-jmx'
|
- job_name: 'kubernetes-pod-jmx'
|
||||||
sample_limit: {{ .Values.scrapeSampleLimit }}
|
sample_limit: {{ .Values.scrapeSampleLimit }}
|
||||||
metrics_path: /{{ .Values.prometheusMetricsEndpoint }}
|
|
||||||
kubernetes_sd_configs:
|
kubernetes_sd_configs:
|
||||||
- role: pod
|
- role: pod
|
||||||
relabel_configs:
|
relabel_configs:
|
||||||
@@ -46,22 +48,24 @@ spec:
|
|||||||
regex: 'jvm_gc_collection_seconds.*'
|
regex: 'jvm_gc_collection_seconds.*'
|
||||||
action: drop
|
action: drop
|
||||||
exporters:
|
exporters:
|
||||||
awsprometheusremotewrite:
|
prometheusremotewrite:
|
||||||
endpoint: {{ .Values.ampurl }}
|
endpoint: {{ .Values.ampurl }}
|
||||||
aws_auth:
|
auth:
|
||||||
region: {{ .Values.region }}
|
authenticator: sigv4auth
|
||||||
service: "aps"
|
|
||||||
logging:
|
logging:
|
||||||
loglevel: info
|
loglevel: info
|
||||||
extensions:
|
extensions:
|
||||||
|
sigv4auth:
|
||||||
|
region: {{ .Values.region }}
|
||||||
|
service: "aps"
|
||||||
health_check:
|
health_check:
|
||||||
pprof:
|
pprof:
|
||||||
endpoint: :1888
|
endpoint: :1888
|
||||||
zpages:
|
zpages:
|
||||||
endpoint: :55679
|
endpoint: :55679
|
||||||
service:
|
service:
|
||||||
extensions: [pprof, zpages, health_check]
|
extensions: [pprof, zpages, health_check, sigv4auth]
|
||||||
pipelines:
|
pipelines:
|
||||||
metrics:
|
metrics:
|
||||||
receivers: [prometheus]
|
receivers: [prometheus]
|
||||||
exporters: [logging, awsprometheusremotewrite]
|
exporters: [logging, prometheusremotewrite]
|
||||||
|
|||||||
@@ -1,7 +1,6 @@
|
|||||||
ampurl: ${amp_url}
|
ampurl: ${amp_url}
|
||||||
region: ${region}
|
region: ${region}
|
||||||
prometheusMetricsEndpoint: ${prometheus_metrics_endpoint}
|
prometheusMetricsEndpoint: ${prometheus_metrics_endpoint}
|
||||||
prometheusMetricsPort: ${prometheus_metrics_port}
|
globalScrapeInterval: ${scrape_interval}
|
||||||
scrapeInterval: ${scrape_interval}
|
globalScrapeTimeout: ${scrape_timeout}
|
||||||
scrapeTimeout: ${scrape_timeout}
|
|
||||||
scrapeSampleLimit: ${scrape_sample_limit}
|
scrapeSampleLimit: ${scrape_sample_limit}
|
||||||
|
|||||||
@@ -0,0 +1,6 @@
|
|||||||
|
output "grafana_dashboard_urls" {
|
||||||
|
value = [concat(
|
||||||
|
grafana_dashboard.this.*.url,
|
||||||
|
)]
|
||||||
|
description = "URLs for dashboards created"
|
||||||
|
}
|
||||||
|
|||||||
@@ -1,17 +1,48 @@
|
|||||||
|
variable "eks_cluster_id" {
|
||||||
|
description = "EKS Cluster Id"
|
||||||
|
type = string
|
||||||
|
}
|
||||||
|
|
||||||
|
variable "irsa_iam_role_path" {
|
||||||
|
description = "IAM role path for IRSA roles"
|
||||||
|
type = string
|
||||||
|
default = "/"
|
||||||
|
}
|
||||||
|
|
||||||
|
variable "irsa_iam_permissions_boundary" {
|
||||||
|
description = "IAM permissions boundary for IRSA roles"
|
||||||
|
type = string
|
||||||
|
default = ""
|
||||||
|
}
|
||||||
|
|
||||||
variable "enable_recording_rules" {
|
variable "enable_recording_rules" {
|
||||||
description = "Enable AMP recording rules"
|
description = "Enables or disables Managed Prometheus recording rules. Disabling this might affect some data in the dashboards"
|
||||||
type = bool
|
type = bool
|
||||||
default = true
|
default = true
|
||||||
}
|
}
|
||||||
|
|
||||||
variable "amp_endpoint" {
|
variable "enable_alerting_rules" {
|
||||||
description = "Amazon Managed Prometheus endpoint"
|
description = "Enables or disables Managed Prometheus alerting rules"
|
||||||
type = string
|
type = bool
|
||||||
|
default = true
|
||||||
}
|
}
|
||||||
|
|
||||||
variable "amp_id" {
|
variable "managed_prometheus_workspace_endpoint" {
|
||||||
description = "Managed Prometheus workspace id"
|
description = "Amazon Managed Prometheus Workspace Endpoint"
|
||||||
type = string
|
type = string
|
||||||
|
default = null
|
||||||
|
}
|
||||||
|
|
||||||
|
variable "managed_prometheus_workspace_id" {
|
||||||
|
description = "Amazon Managed Prometheus Workspace ID"
|
||||||
|
type = string
|
||||||
|
default = null
|
||||||
|
}
|
||||||
|
|
||||||
|
variable "managed_prometheus_workspace_region" {
|
||||||
|
description = "Amazon Managed Prometheus Workspace's Region"
|
||||||
|
type = string
|
||||||
|
default = null
|
||||||
}
|
}
|
||||||
|
|
||||||
variable "helm_config" {
|
variable "helm_config" {
|
||||||
@@ -20,30 +51,29 @@ variable "helm_config" {
|
|||||||
default = {}
|
default = {}
|
||||||
}
|
}
|
||||||
|
|
||||||
variable "amp_region" {
|
|
||||||
description = "Amazon Managed Prometheus Workspace's Region"
|
|
||||||
type = string
|
|
||||||
default = null
|
|
||||||
}
|
|
||||||
|
|
||||||
variable "dashboards_folder_id" {
|
variable "dashboards_folder_id" {
|
||||||
description = "Grafana folder ID for automatic dashboards"
|
description = "Grafana folder ID for automatic dashboards"
|
||||||
type = string
|
type = string
|
||||||
}
|
}
|
||||||
|
|
||||||
variable "addon_context" {
|
variable "prometheus_config" {
|
||||||
description = "Input configuration for the addon"
|
description = "Controls default values such as scrape interval, timeouts and ports globally"
|
||||||
type = object({
|
type = object({
|
||||||
aws_caller_identity_account_id = string
|
global_scrape_interval = string
|
||||||
aws_caller_identity_arn = string
|
global_scrape_timeout = string
|
||||||
aws_eks_cluster_endpoint = string
|
scrape_sample_limit = number
|
||||||
aws_partition_id = string
|
|
||||||
aws_region_name = string
|
|
||||||
eks_cluster_id = string
|
|
||||||
eks_oidc_issuer_url = string
|
|
||||||
eks_oidc_provider_arn = string
|
|
||||||
irsa_iam_permissions_boundary = string
|
|
||||||
irsa_iam_role_path = string
|
|
||||||
tags = map(string)
|
|
||||||
})
|
})
|
||||||
|
|
||||||
|
default = {
|
||||||
|
global_scrape_interval = "60s"
|
||||||
|
global_scrape_timeout = "15s"
|
||||||
|
scrape_sample_limit = 1000
|
||||||
|
}
|
||||||
|
nullable = false
|
||||||
|
}
|
||||||
|
|
||||||
|
variable "tags" {
|
||||||
|
description = "Additional tags (e.g. `map('BusinessUnit`,`XYZ`)"
|
||||||
|
type = map(string)
|
||||||
|
default = {}
|
||||||
}
|
}
|
||||||
|
|||||||
Reference in New Issue
Block a user