Java dev (#52)

* Update Java module

* Add java example

* Updating Readme

* Updating Readme Images

* Pre-commit

* Fixing readme screenshot

* Update README.md

Co-authored-by: Rodrigue Koffi <bonclay7@users.noreply.github.com>
This commit is contained in:
Kevin Lewin
2022-10-14 13:12:14 -04:00
committed by GitHub
parent 4fdb719f32
commit a5a444ee65
12 changed files with 582 additions and 72 deletions
+240
View File
@@ -0,0 +1,240 @@
# Existing Cluster with the AWS Observability accelerator base module and Java monitoring
This example demonstrates how to use the AWS Observability Accelerator Terraform
modules with Java monitoring enabled.
The current example deploys the [AWS Distro for OpenTelemetry Operator](https://docs.aws.amazon.com/eks/latest/userguide/opentelemetry.html) for Amazon EKS with its requirements and make use of existing
Amazon Managed Service for Prometheus and Amazon Managed Grafana workspaces.
It is based on the `java module`, one of our [workloads modules](../../modules/workloads/)
to provide an existing EKS cluster with an OpenTelemetry collector,
curated Grafana dashboards, Prometheus alerting and recording rules with multiple
configuration options on the cluster infrastructure.
## Prerequisites
Ensure that you have the following tools installed locally:
1. [aws cli v2](https://docs.aws.amazon.com/cli/latest/userguide/getting-started-install.html)
2. [kubectl](https://kubernetes.io/docs/tasks/tools/)
3. [terraform](https://learn.hashicorp.com/tutorials/terraform/install-cli)
## Setup
This example uses a local terraform state. If you need states to be saved remotely,
on Amazon S3 for example, visit the [terraform remote states](https://www.terraform.io/language/state/remote) documentation
1. Clone the repo using the command below
```
git clone https://github.com/aws-observability/terraform-aws-observability-accelerator.git
```
2. Initialize terraform
```console
cd examples/existing-cluster-java
terraform init
```
3. AWS Region
Specify the AWS Region where the resources will be deployed. Edit the `terraform.tfvars` file and modify `aws_region="..."`. You can also use environement variables `export TF_VAR_aws_region=xxx`.
4. Amazon EKS Cluster
To run this example, you need to provide your EKS cluster name.
If you don't have a cluster ready, visit [this example](../eks-cluster-with-vpc)
first to create a new one.
Add your cluster name for `eks_cluster_id="..."` to the `terraform.tfvars` or use an environment variable `export TF_VAR_eks_cluster_id=xxx`.
5. Amazon Managed Service for Prometheus workspace (optional)
If you have an existing workspace, add `managed_prometheus_workspace_id=ws-xxx`
or use an environment variable `export TF_VAR_managed_prometheus_workspace_id=ws-xxx`.
If you don't specify anything a new workspace will be created for you.
6. Amazon Managed Grafana workspace
If you have an existing workspace, create an environment variable `export TF_VAR_managed_grafana_workspace_id=g-xxx`.
7. <a name="apikey"></a> Grafana API Key
Amazon Managed Service for Grafana provides a control plane API for generating Grafana API keys. We will provide to Terraform
a short lived API key to run the `apply` or `destroy` command.
Ensure you have necessary IAM permissions (`CreateWorkspaceApiKey, DeleteWorkspaceApiKey`)
```sh
export TF_VAR_grafana_api_key=`aws grafana create-workspace-api-key --key-name "observability-accelerator-$(date +%s)" --key-role ADMIN --seconds-to-live 1200 --workspace-id $TF_VAR_managed_grafana_workspace_id --query key --output text`
```
## Deploy
```sh
terraform apply -var-file=terraform.tfvars
```
or if you had only setup environment variables, run
```sh
terraform apply
```
## Visualization
1. Prometheus datasource on Grafana
Open your Grafana workspace and under Configuration -> Data sources, you will see `aws-observability-accelerator`. Open and click `Save & test`. You will then see a notification confirming that the Amazon Managed Service for Prometheus workspace is ready to be used on Grafana.
2. Grafana dashboards
Go to the Dashboards panel of your Grafana workspace. There will be a folder called `Observability Accelerator Dashboards`
<img width="832" alt="image" src="https://user-images.githubusercontent.com/97046295/194903648-57c55d30-6f90-4b03-9eb6-577aaba7dc22.png">
Open the "Java/JMX" dashboard to view its visualization
![image](https://user-images.githubusercontent.com/10175027/195903211-c47a5746-daa7-41f2-a6ea-bfe13f630c63.png)
2. Amazon Managed Service for Prometheus rules and alerts
Open the Amazon Managed Service for Prometheus console and view the details of your workspace. Under the `Rules management` tab, you will find new rules deployed.
<img width="1314" alt="image" src="https://user-images.githubusercontent.com/97046295/194904104-09a28577-d149-478e-b0a1-dc21cb7effc1.png">
To setup your alert receiver, with Amazon SNS, follow [this documentation](https://docs.aws.amazon.com/prometheus/latest/userguide/AMP-alertmanager-receiver.html)
## Deploy an Example Java Application
In this section we will reuse an example from the AWS OpenTelemetry collector [repository](https://github.com/aws-observability/aws-otel-collector/blob/main/docs/developers/container-insights-eks-jmx.md). For convenience, the steps can be found below.
1. Clone [this repository](https://github.com/aws-observability/aws-otel-test-framework) and navigate to the `sample-apps/jmx/` directory.
2. Authenticate to Amazon ECR
```sh
export AWS_ACCOUNT_ID=`aws sts get-caller-identity --query Account --output text`
export AWS_REGION={region}
aws ecr get-login-password --region $AWS_REGION | docker login --username AWS --password-stdin $AWS_ACCOUNT_ID.dkr.ecr.$AWS_REGION.amazonaws.com
```
3. Create an Amazon ECR repository
```sh
aws ecr create-repository --repository-name prometheus-sample-tomcat-jmx \
--image-scanning-configuration scanOnPush=true \
--region $AWS_REGION
```
4. Build Docker image and push to ECR.
```sh
docker build -t $AWS_ACCOUNT_ID.dkr.ecr.$AWS_REGION.amazonaws.com/prometheus-sample-tomcat-jmx:latest .
docker push $AWS_ACCOUNT_ID.dkr.ecr.$AWS_REGION.amazonaws.com/prometheus-sample-tomcat-jmx:latest
```
5. Install sample application
```sh
export SAMPLE_TRAFFIC_NAMESPACE=javajmx-sample
curl https://raw.githubusercontent.com/aws-observability/aws-otel-test-framework/terraform/sample-apps/jmx/examples/prometheus-metrics-sample.yaml > metrics-sample.yaml
sed -i "s/{{aws_account_id}}/$AWS_ACCOUNT_ID/g" metrics-sample.yaml
sed -i "s/{{region}}/$AWS_REGION/g" metrics-sample.yaml
sed -i "s/{{namespace}}/$SAMPLE_TRAFFIC_NAMESPACE/g" metrics-sample.yaml
kubectl apply -f metrics-sample.yaml
```
Verify that the sample application is running:
```sh
kubectl get pods -n $SAMPLE_TRAFFIC_NAMESPACE
NAME READY STATUS RESTARTS AGE
tomcat-bad-traffic-generator 1/1 Running 0 11s
tomcat-example-7958666589-2q755 0/1 ContainerCreating 0 11s
tomcat-traffic-generator 1/1 Running 0 11s
```
## Advanced configuration
1. Cross-region Amazon Managed Prometheus workspace
If your existing Amazon Managed Prometheus workspace is in another AWS Region,
add this `managed_prometheus_region=xxx` and `managed_prometheus_workspace_id=ws-xxx`.
2. Cross-region Amazon Managed Grafana workspace
If your existing Amazon Managed Prometheus workspace is in another AWS Region,
add this `managed_prometheus_region=xxx` and `managed_prometheus_workspace_id=ws-xxx`.
## Destroy resources
If you leave this stack running, you will continue to incur charges. To remove all resources
created by Terraform, [refresh your Grafana API key](#apikey) and run:
```sh
terraform destroy -var-file=terraform.tfvars
```
<!-- BEGINNING OF PRE-COMMIT-TERRAFORM DOCS HOOK -->
## Requirements
| Name | Version |
|------|---------|
| <a name="requirement_terraform"></a> [terraform](#requirement\_terraform) | >= 1.1.0, < 1.3.0 |
| <a name="requirement_aws"></a> [aws](#requirement\_aws) | >= 4.0.0 |
| <a name="requirement_grafana"></a> [grafana](#requirement\_grafana) | >= 1.25.0 |
| <a name="requirement_helm"></a> [helm](#requirement\_helm) | >= 2.4.1 |
| <a name="requirement_kubectl"></a> [kubectl](#requirement\_kubectl) | >= 1.14 |
| <a name="requirement_kubernetes"></a> [kubernetes](#requirement\_kubernetes) | >= 2.10 |
## Providers
| Name | Version |
|------|---------|
| <a name="provider_aws"></a> [aws](#provider\_aws) | >= 4.0.0 |
## Modules
| Name | Source | Version |
|------|--------|---------|
| <a name="module_eks_observability_accelerator"></a> [eks\_observability\_accelerator](#module\_eks\_observability\_accelerator) | ../../ | n/a |
| <a name="module_workloads_java"></a> [workloads\_java](#module\_workloads\_java) | ../../modules/workloads/java | n/a |
## Resources
| Name | Type |
|------|------|
| [aws_eks_cluster.this](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/data-sources/eks_cluster) | data source |
| [aws_eks_cluster_auth.this](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/data-sources/eks_cluster_auth) | data source |
## Inputs
| Name | Description | Type | Default | Required |
|------|-------------|------|---------|:--------:|
| <a name="input_aws_region"></a> [aws\_region](#input\_aws\_region) | AWS Region | `string` | n/a | yes |
| <a name="input_eks_cluster_id"></a> [eks\_cluster\_id](#input\_eks\_cluster\_id) | Name of the EKS cluster | `string` | n/a | yes |
| <a name="input_grafana_api_key"></a> [grafana\_api\_key](#input\_grafana\_api\_key) | API key for authorizing the Grafana provider to make changes to Amazon Managed Grafana | `string` | `""` | no |
| <a name="input_managed_grafana_workspace_id"></a> [managed\_grafana\_workspace\_id](#input\_managed\_grafana\_workspace\_id) | Amazon Managed Grafana Workspace ID | `string` | `""` | no |
| <a name="input_managed_prometheus_workspace_id"></a> [managed\_prometheus\_workspace\_id](#input\_managed\_prometheus\_workspace\_id) | Amazon Managed Service for Prometheus Workspace ID | `string` | `""` | no |
## Outputs
| Name | Description |
|------|-------------|
| <a name="output_aws_region"></a> [aws\_region](#output\_aws\_region) | AWS Region |
| <a name="output_eks_cluster_id"></a> [eks\_cluster\_id](#output\_eks\_cluster\_id) | EKS Cluster Id |
| <a name="output_eks_cluster_version"></a> [eks\_cluster\_version](#output\_eks\_cluster\_version) | EKS Cluster version |
| <a name="output_grafana_dashboard_urls"></a> [grafana\_dashboard\_urls](#output\_grafana\_dashboard\_urls) | URLs for dashboards created |
| <a name="output_managed_prometheus_workspace_endpoint"></a> [managed\_prometheus\_workspace\_endpoint](#output\_managed\_prometheus\_workspace\_endpoint) | Amazon Managed Prometheus workspace endpoint |
| <a name="output_managed_prometheus_workspace_id"></a> [managed\_prometheus\_workspace\_id](#output\_managed\_prometheus\_workspace\_id) | Amazon Managed Prometheus workspace ID |
<!-- END OF PRE-COMMIT-TERRAFORM DOCS HOOK -->
+100
View File
@@ -0,0 +1,100 @@
provider "aws" {
region = local.region
}
data "aws_eks_cluster_auth" "this" {
name = var.eks_cluster_id
}
data "aws_eks_cluster" "this" {
name = var.eks_cluster_id
}
provider "kubernetes" {
host = local.eks_cluster_endpoint
cluster_ca_certificate = base64decode(data.aws_eks_cluster.this.certificate_authority[0].data)
token = data.aws_eks_cluster_auth.this.token
}
provider "helm" {
kubernetes {
host = local.eks_cluster_endpoint
cluster_ca_certificate = base64decode(data.aws_eks_cluster.this.certificate_authority[0].data)
token = data.aws_eks_cluster_auth.this.token
}
}
locals {
region = var.aws_region
eks_cluster_endpoint = data.aws_eks_cluster.this.endpoint
create_new_workspace = var.managed_prometheus_workspace_id == "" ? true : false
tags = {
Source = "github.com/aws-observability/terraform-aws-observability-accelerator"
}
}
# deploys the base module
module "eks_observability_accelerator" {
# source = "aws-observability/terrarom-aws-observability-accelerator"
source = "../../"
aws_region = var.aws_region
eks_cluster_id = var.eks_cluster_id
# deploys AWS Distro for OpenTelemetry operator into the cluster
enable_amazon_eks_adot = true
# reusing existing certificate manager? defaults to true
enable_cert_manager = true
# creates a new Amazon Managed Prometheus workspace, defaults to true
enable_managed_prometheus = local.create_new_workspace
# reusing existing Amazon Managed Prometheus if specified
managed_prometheus_workspace_id = var.managed_prometheus_workspace_id
managed_prometheus_workspace_region = null # defaults to the current region, useful for cross region scenarios (same account)
# sets up the Amazon Managed Prometheus alert manager at the workspace level
enable_alertmanager = true
# reusing existing Amazon Managed Grafana workspace
enable_managed_grafana = false
managed_grafana_workspace_id = var.managed_grafana_workspace_id
grafana_api_key = var.grafana_api_key
tags = local.tags
}
# https://www.terraform.io/language/modules/develop/providers
# A module intended to be called by one or more other modules must not contain
# any provider blocks.
# This allows forcing dependency between base and workloads module
provider "grafana" {
url = module.eks_observability_accelerator.managed_grafana_workspace_endpoint
auth = var.grafana_api_key
}
module "workloads_java" {
source = "../../modules/workloads/java"
eks_cluster_id = module.eks_observability_accelerator.eks_cluster_id
dashboards_folder_id = module.eks_observability_accelerator.grafana_dashboards_folder_id
managed_prometheus_workspace_id = module.eks_observability_accelerator.managed_prometheus_workspace_id
managed_prometheus_workspace_endpoint = module.eks_observability_accelerator.managed_prometheus_workspace_endpoint
managed_prometheus_workspace_region = module.eks_observability_accelerator.managed_prometheus_workspace_region
# optional, defaults to 60s interval and 15s timeout
prometheus_config = {
global_scrape_interval = "60s"
global_scrape_timeout = "15s"
scrape_sample_limit = 2000
}
tags = local.tags
depends_on = [
module.eks_observability_accelerator
]
}
+29
View File
@@ -0,0 +1,29 @@
output "eks_cluster_id" {
description = "EKS Cluster Id"
value = module.eks_observability_accelerator.eks_cluster_id
}
output "aws_region" {
description = "AWS Region"
value = module.eks_observability_accelerator.aws_region
}
output "eks_cluster_version" {
description = "EKS Cluster version"
value = module.eks_observability_accelerator.eks_cluster_version
}
output "managed_prometheus_workspace_endpoint" {
description = "Amazon Managed Prometheus workspace endpoint"
value = module.eks_observability_accelerator.managed_prometheus_workspace_endpoint
}
output "managed_prometheus_workspace_id" {
description = "Amazon Managed Prometheus workspace ID"
value = module.eks_observability_accelerator.managed_prometheus_workspace_id
}
output "grafana_dashboard_urls" {
description = "URLs for dashboards created"
value = module.workloads_java.grafana_dashboard_urls
}
@@ -0,0 +1,24 @@
variable "eks_cluster_id" {
description = "Name of the EKS cluster"
type = string
}
variable "aws_region" {
description = "AWS Region"
type = string
}
variable "managed_prometheus_workspace_id" {
description = "Amazon Managed Service for Prometheus Workspace ID"
type = string
default = ""
}
variable "managed_grafana_workspace_id" {
description = "Amazon Managed Grafana Workspace ID"
type = string
default = ""
}
variable "grafana_api_key" {
description = "API key for authorizing the Grafana provider to make changes to Amazon Managed Grafana"
type = string
default = ""
sensitive = true
}
@@ -0,0 +1,34 @@
terraform {
required_version = ">= 1.1.0, < 1.3.0"
required_providers {
aws = {
source = "hashicorp/aws"
version = ">= 4.0.0"
}
kubernetes = {
source = "hashicorp/kubernetes"
version = ">= 2.10"
}
kubectl = {
source = "gavinbunney/kubectl"
version = ">= 1.14"
}
helm = {
source = "hashicorp/helm"
version = ">= 2.4.1"
}
grafana = {
source = "grafana/grafana"
version = ">= 1.25.0"
}
}
# ## Used for end-to-end testing on project; update to suit your needs
# backend "s3" {
# bucket = "observability-accelerator-terraform-states"
# region = "us-west-2"
# key = "e2e/existing-cluster-with-base-and-infra/terraform.tfstate"
# }
}
+19 -8
View File
@@ -29,29 +29,40 @@ This module provides monitoring for Java based workloads with the following reso
| Name | Source | Version | | Name | Source | Version |
|------|--------|---------| |------|--------|---------|
| <a name="module_helm_addon"></a> [helm\_addon](#module\_helm\_addon) | github.com/aws-ia/terraform-aws-eks-blueprints//modules/kubernetes-addons/helm-addon | v4.8.1 | | <a name="module_helm_addon"></a> [helm\_addon](#module\_helm\_addon) | github.com/aws-ia/terraform-aws-eks-blueprints//modules/kubernetes-addons/helm-addon | v4.12.1 |
## Resources ## Resources
| Name | Type | | Name | Type |
|------|------| |------|------|
| [aws_prometheus_rule_group_namespace.this](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/resources/prometheus_rule_group_namespace) | resource | | [aws_prometheus_rule_group_namespace.alerting_rules](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/resources/prometheus_rule_group_namespace) | resource |
| [aws_prometheus_rule_group_namespace.recording_rules](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/resources/prometheus_rule_group_namespace) | resource |
| [grafana_dashboard.this](https://registry.terraform.io/providers/grafana/grafana/latest/docs/resources/dashboard) | resource | | [grafana_dashboard.this](https://registry.terraform.io/providers/grafana/grafana/latest/docs/resources/dashboard) | resource |
| [aws_caller_identity.current](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/data-sources/caller_identity) | data source |
| [aws_eks_cluster.eks_cluster](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/data-sources/eks_cluster) | data source |
| [aws_partition.current](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/data-sources/partition) | data source | | [aws_partition.current](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/data-sources/partition) | data source |
| [aws_region.current](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/data-sources/region) | data source |
## Inputs ## Inputs
| Name | Description | Type | Default | Required | | Name | Description | Type | Default | Required |
|------|-------------|------|---------|:--------:| |------|-------------|------|---------|:--------:|
| <a name="input_addon_context"></a> [addon\_context](#input\_addon\_context) | Input configuration for the addon | <pre>object({<br> aws_caller_identity_account_id = string<br> aws_caller_identity_arn = string<br> aws_eks_cluster_endpoint = string<br> aws_partition_id = string<br> aws_region_name = string<br> eks_cluster_id = string<br> eks_oidc_issuer_url = string<br> eks_oidc_provider_arn = string<br> irsa_iam_permissions_boundary = string<br> irsa_iam_role_path = string<br> tags = map(string)<br> })</pre> | n/a | yes |
| <a name="input_amp_endpoint"></a> [amp\_endpoint](#input\_amp\_endpoint) | Amazon Managed Prometheus endpoint | `string` | n/a | yes |
| <a name="input_amp_id"></a> [amp\_id](#input\_amp\_id) | Managed Prometheus workspace id | `string` | n/a | yes |
| <a name="input_amp_region"></a> [amp\_region](#input\_amp\_region) | Amazon Managed Prometheus Workspace's Region | `string` | `null` | no |
| <a name="input_dashboards_folder_id"></a> [dashboards\_folder\_id](#input\_dashboards\_folder\_id) | Grafana folder ID for automatic dashboards | `string` | n/a | yes | | <a name="input_dashboards_folder_id"></a> [dashboards\_folder\_id](#input\_dashboards\_folder\_id) | Grafana folder ID for automatic dashboards | `string` | n/a | yes |
| <a name="input_enable_recording_rules"></a> [enable\_recording\_rules](#input\_enable\_recording\_rules) | Enable AMP recording rules | `bool` | `true` | no | | <a name="input_eks_cluster_id"></a> [eks\_cluster\_id](#input\_eks\_cluster\_id) | EKS Cluster Id | `string` | n/a | yes |
| <a name="input_enable_alerting_rules"></a> [enable\_alerting\_rules](#input\_enable\_alerting\_rules) | Enables or disables Managed Prometheus alerting rules | `bool` | `true` | no |
| <a name="input_enable_recording_rules"></a> [enable\_recording\_rules](#input\_enable\_recording\_rules) | Enables or disables Managed Prometheus recording rules. Disabling this might affect some data in the dashboards | `bool` | `true` | no |
| <a name="input_helm_config"></a> [helm\_config](#input\_helm\_config) | Helm Config for Prometheus | `any` | `{}` | no | | <a name="input_helm_config"></a> [helm\_config](#input\_helm\_config) | Helm Config for Prometheus | `any` | `{}` | no |
| <a name="input_irsa_iam_permissions_boundary"></a> [irsa\_iam\_permissions\_boundary](#input\_irsa\_iam\_permissions\_boundary) | IAM permissions boundary for IRSA roles | `string` | `""` | no |
| <a name="input_irsa_iam_role_path"></a> [irsa\_iam\_role\_path](#input\_irsa\_iam\_role\_path) | IAM role path for IRSA roles | `string` | `"/"` | no |
| <a name="input_managed_prometheus_workspace_endpoint"></a> [managed\_prometheus\_workspace\_endpoint](#input\_managed\_prometheus\_workspace\_endpoint) | Amazon Managed Prometheus Workspace Endpoint | `string` | `null` | no |
| <a name="input_managed_prometheus_workspace_id"></a> [managed\_prometheus\_workspace\_id](#input\_managed\_prometheus\_workspace\_id) | Amazon Managed Prometheus Workspace ID | `string` | `null` | no |
| <a name="input_managed_prometheus_workspace_region"></a> [managed\_prometheus\_workspace\_region](#input\_managed\_prometheus\_workspace\_region) | Amazon Managed Prometheus Workspace's Region | `string` | `null` | no |
| <a name="input_prometheus_config"></a> [prometheus\_config](#input\_prometheus\_config) | Controls default values such as scrape interval, timeouts and ports globally | <pre>object({<br> global_scrape_interval = string<br> global_scrape_timeout = string<br> scrape_sample_limit = number<br> })</pre> | <pre>{<br> "global_scrape_interval": "60s",<br> "global_scrape_timeout": "15s",<br> "scrape_sample_limit": 1000<br>}</pre> | no |
| <a name="input_tags"></a> [tags](#input\_tags) | Additional tags (e.g. `map('BusinessUnit`,`XYZ`) | `map(string)` | `{}` | no |
## Outputs ## Outputs
No outputs. | Name | Description |
|------|-------------|
| <a name="output_grafana_dashboard_urls"></a> [grafana\_dashboard\_urls](#output\_grafana\_dashboard\_urls) | URLs for dashboards created |
<!-- END OF PRE-COMMIT-TERRAFORM DOCS HOOK --> <!-- END OF PRE-COMMIT-TERRAFORM DOCS HOOK -->
+31
View File
@@ -0,0 +1,31 @@
data "aws_partition" "current" {}
data "aws_caller_identity" "current" {}
data "aws_region" "current" {}
data "aws_eks_cluster" "eks_cluster" {
name = var.eks_cluster_id
}
locals {
name = "adot-collector-java"
namespace = try(var.helm_config.namespace, local.name)
eks_oidc_issuer_url = replace(data.aws_eks_cluster.eks_cluster.identity[0].oidc[0].issuer, "https://", "")
eks_cluster_endpoint = data.aws_eks_cluster.eks_cluster.endpoint
context = {
aws_caller_identity_account_id = data.aws_caller_identity.current.account_id
aws_caller_identity_arn = data.aws_caller_identity.current.arn
aws_eks_cluster_endpoint = local.eks_cluster_endpoint
aws_partition_id = data.aws_partition.current.partition
aws_region_name = data.aws_region.current.name
eks_cluster_id = var.eks_cluster_id
eks_oidc_issuer_url = local.eks_oidc_issuer_url
eks_oidc_provider_arn = "arn:${data.aws_partition.current.partition}:iam::${data.aws_caller_identity.current.account_id}:oidc-provider/${local.eks_oidc_issuer_url}"
tags = var.tags
irsa_iam_role_path = var.irsa_iam_role_path
irsa_iam_permissions_boundary = var.irsa_iam_permissions_boundary
}
}
+26 -24
View File
@@ -1,13 +1,6 @@
locals {
name = "adot-collector-java"
namespace = try(var.helm_config.namespace, local.name)
}
data "aws_partition" "current" {}
# deploys collector # deploys collector
module "helm_addon" { module "helm_addon" {
source = "github.com/aws-ia/terraform-aws-eks-blueprints//modules/kubernetes-addons/helm-addon?ref=v4.8.1" source = "github.com/aws-ia/terraform-aws-eks-blueprints//modules/kubernetes-addons/helm-addon?ref=v4.12.1"
helm_config = merge( helm_config = merge(
{ {
@@ -23,31 +16,31 @@ module "helm_addon" {
set_values = [ set_values = [
{ {
name = "ampurl" name = "ampurl"
value = "${var.amp_endpoint}api/v1/remote_write" value = "${var.managed_prometheus_workspace_endpoint}api/v1/remote_write"
}, },
{ {
name = "region" name = "region"
value = var.amp_region value = var.managed_prometheus_workspace_region
}, },
{ {
name = "prometheusMetricsEndpoint" name = "ekscluster"
value = "metrics" value = local.context.eks_cluster_id
}, },
{ {
name = "prometheusMetricsPort" name = "accountId"
value = 8888 value = local.context.aws_caller_identity_account_id
}, },
{ {
name = "scrapeInterval" name = "globalScrapeInterval"
value = "15s" value = var.prometheus_config.global_scrape_interval
}, },
{ {
name = "scrapeTimeout" name = "globalScrapeTimeout"
value = "10s" value = var.prometheus_config.global_scrape_timeout
}, },
{ {
name = "scrapeSampleLimit" name = "scrapeSampleLimit"
value = 1000 value = var.prometheus_config.scrape_sample_limit
} }
] ]
@@ -59,21 +52,30 @@ module "helm_addon" {
irsa_iam_policies = ["arn:${data.aws_partition.current.partition}:iam::aws:policy/AmazonPrometheusRemoteWriteAccess"] irsa_iam_policies = ["arn:${data.aws_partition.current.partition}:iam::aws:policy/AmazonPrometheusRemoteWriteAccess"]
} }
addon_context = var.addon_context addon_context = local.context
} }
resource "aws_prometheus_rule_group_namespace" "recording_rules" {
resource "aws_prometheus_rule_group_namespace" "this" {
count = var.enable_recording_rules ? 1 : 0 count = var.enable_recording_rules ? 1 : 0
name = "java_rules" name = "accelerator-java-rules"
workspace_id = var.amp_id workspace_id = var.managed_prometheus_workspace_id
data = <<EOF data = <<EOF
groups: groups:
- name: default-metric - name: default-metric
rules: rules:
- record: metric:recording_rule - record: metric:recording_rule
expr: avg(rate(container_cpu_usage_seconds_total[5m])) expr: avg(rate(container_cpu_usage_seconds_total[5m]))
EOF
}
resource "aws_prometheus_rule_group_namespace" "alerting_rules" {
count = var.enable_alerting_rules ? 1 : 0
name = "accelerator-java-alerting"
workspace_id = var.managed_prometheus_workspace_id
data = <<EOF
groups:
- name: default-alert - name: default-alert
rules: rules:
- alert: metric:alerting_rule - alert: metric:alerting_rule
@@ -3,7 +3,7 @@ kind: OpenTelemetryCollector
metadata: metadata:
name: adot name: adot
spec: spec:
image: public.ecr.aws/aws-observability/aws-otel-collector:latest image: public.ecr.aws/aws-observability/aws-otel-collector:v0.22.0
mode: deployment mode: deployment
serviceAccount: adot-collector-java serviceAccount: adot-collector-java
config: | config: |
@@ -11,13 +11,15 @@ spec:
prometheus: prometheus:
config: config:
global: global:
scrape_interval: {{ .Values.scrapeInterval }} scrape_interval: {{ .Values.globalScrapeInterval }}
scrape_timeout: {{ .Values.scrapeTimeout }} scrape_timeout: {{ .Values.globalScrapeTimeout }}
external_labels:
cluster: {{ .Values.ekscluster }}
account_id: {{ .Values.accountId }}
region: {{ .Values.region }}
scrape_configs: scrape_configs:
- job_name: 'kubernetes-pod-jmx' - job_name: 'kubernetes-pod-jmx'
sample_limit: {{ .Values.scrapeSampleLimit }} sample_limit: {{ .Values.scrapeSampleLimit }}
metrics_path: /{{ .Values.prometheusMetricsEndpoint }}
kubernetes_sd_configs: kubernetes_sd_configs:
- role: pod - role: pod
relabel_configs: relabel_configs:
@@ -46,22 +48,24 @@ spec:
regex: 'jvm_gc_collection_seconds.*' regex: 'jvm_gc_collection_seconds.*'
action: drop action: drop
exporters: exporters:
awsprometheusremotewrite: prometheusremotewrite:
endpoint: {{ .Values.ampurl }} endpoint: {{ .Values.ampurl }}
aws_auth: auth:
region: {{ .Values.region }} authenticator: sigv4auth
service: "aps"
logging: logging:
loglevel: info loglevel: info
extensions: extensions:
sigv4auth:
region: {{ .Values.region }}
service: "aps"
health_check: health_check:
pprof: pprof:
endpoint: :1888 endpoint: :1888
zpages: zpages:
endpoint: :55679 endpoint: :55679
service: service:
extensions: [pprof, zpages, health_check] extensions: [pprof, zpages, health_check, sigv4auth]
pipelines: pipelines:
metrics: metrics:
receivers: [prometheus] receivers: [prometheus]
exporters: [logging, awsprometheusremotewrite] exporters: [logging, prometheusremotewrite]
@@ -1,7 +1,6 @@
ampurl: ${amp_url} ampurl: ${amp_url}
region: ${region} region: ${region}
prometheusMetricsEndpoint: ${prometheus_metrics_endpoint} prometheusMetricsEndpoint: ${prometheus_metrics_endpoint}
prometheusMetricsPort: ${prometheus_metrics_port} globalScrapeInterval: ${scrape_interval}
scrapeInterval: ${scrape_interval} globalScrapeTimeout: ${scrape_timeout}
scrapeTimeout: ${scrape_timeout}
scrapeSampleLimit: ${scrape_sample_limit} scrapeSampleLimit: ${scrape_sample_limit}
+6
View File
@@ -0,0 +1,6 @@
output "grafana_dashboard_urls" {
value = [concat(
grafana_dashboard.this.*.url,
)]
description = "URLs for dashboards created"
}
+55 -25
View File
@@ -1,17 +1,48 @@
variable "eks_cluster_id" {
description = "EKS Cluster Id"
type = string
}
variable "irsa_iam_role_path" {
description = "IAM role path for IRSA roles"
type = string
default = "/"
}
variable "irsa_iam_permissions_boundary" {
description = "IAM permissions boundary for IRSA roles"
type = string
default = ""
}
variable "enable_recording_rules" { variable "enable_recording_rules" {
description = "Enable AMP recording rules" description = "Enables or disables Managed Prometheus recording rules. Disabling this might affect some data in the dashboards"
type = bool type = bool
default = true default = true
} }
variable "amp_endpoint" { variable "enable_alerting_rules" {
description = "Amazon Managed Prometheus endpoint" description = "Enables or disables Managed Prometheus alerting rules"
type = string type = bool
default = true
} }
variable "amp_id" { variable "managed_prometheus_workspace_endpoint" {
description = "Managed Prometheus workspace id" description = "Amazon Managed Prometheus Workspace Endpoint"
type = string type = string
default = null
}
variable "managed_prometheus_workspace_id" {
description = "Amazon Managed Prometheus Workspace ID"
type = string
default = null
}
variable "managed_prometheus_workspace_region" {
description = "Amazon Managed Prometheus Workspace's Region"
type = string
default = null
} }
variable "helm_config" { variable "helm_config" {
@@ -20,30 +51,29 @@ variable "helm_config" {
default = {} default = {}
} }
variable "amp_region" {
description = "Amazon Managed Prometheus Workspace's Region"
type = string
default = null
}
variable "dashboards_folder_id" { variable "dashboards_folder_id" {
description = "Grafana folder ID for automatic dashboards" description = "Grafana folder ID for automatic dashboards"
type = string type = string
} }
variable "addon_context" { variable "prometheus_config" {
description = "Input configuration for the addon" description = "Controls default values such as scrape interval, timeouts and ports globally"
type = object({ type = object({
aws_caller_identity_account_id = string global_scrape_interval = string
aws_caller_identity_arn = string global_scrape_timeout = string
aws_eks_cluster_endpoint = string scrape_sample_limit = number
aws_partition_id = string
aws_region_name = string
eks_cluster_id = string
eks_oidc_issuer_url = string
eks_oidc_provider_arn = string
irsa_iam_permissions_boundary = string
irsa_iam_role_path = string
tags = map(string)
}) })
default = {
global_scrape_interval = "60s"
global_scrape_timeout = "15s"
scrape_sample_limit = 1000
}
nullable = false
}
variable "tags" {
description = "Additional tags (e.g. `map('BusinessUnit`,`XYZ`)"
type = map(string)
default = {}
} }