Adding Module and Example for ECS cluster monitoring with ecs_observer (#211)

* Adding Module and Example for ECS cluster monitoring with ecs_observer

* Adding Module and Example for ECS cluster monitoring with ecs_observer

* Incorporating PR comments

* Restructuring Examples and modules folder for ECS, Added content in main Readme

* Fixing path as per PR comments

* Parameterzing the config files, incorporated PR review comments

* Adding condition for AMP WS and fixing AMP endpoint

* Adding Document for ECS Monitoring and parameterized some variables

* Added sample dashboard

* Adding Document for ECS Monitoring and parameterized some variables

* Fixing failures detected by pre-commit

* Fixing failures detected by pre-commit

* Fixing failures detected by pre-commit

* Pre-commit fixes

* Fixing failures detected by pre-commit

* Fixing failures detected by pre-commit

* Pre-commit

* Fixing HIGH security alerts detected by pre-commit

* Fixing HIGH security alerts detected by pre-commit

* Fixing HIGH security alerts detected by pre-commit, 31stOct

* Add links after merge

* 2ndNov - Added condiotnal creation for Grafana WS and module versions for AMG, AMP

---------

Co-authored-by: Rodrigue Koffi <bonclay7@users.noreply.github.com>
This commit is contained in:
Ruchika Modi
2023-11-02 17:23:32 +05:30
committed by GitHub
parent 0b42935d8b
commit 70405b9732
17 changed files with 863 additions and 0 deletions
+78
View File
@@ -0,0 +1,78 @@
# Observability Module for ECS Monitoring using ecs_observer
This module provides ECS cluster monitoring with the following resources:
- AWS Distro For OpenTelemetry Operator and Collector for Metrics and Traces
- Creates Grafana Dashboards on Amazon Managed Grafana.
- Creates SSM Parameter to store and distribute the ADOT config file
## Pre-requisites
* ECS Cluster with EC2 using examples --> ecs-cluster-with-vpc
* Create a `Prometheus Workspace` either using the Console or using the commented code under modules/ecs-monitoring/main.tf.
* Update your exisitng App(workload) *ECS Task Definition* to add below label/environment variable
- Set ***ECS_PROMETHEUS_EXPORTER_PORT*** to point to the containerPort where the Prometheus metrics are exposed
- Set ***Java_EMF_Metrics*** to true. The CloudWatch agent uses this flag to generated the embedded metric format in the log event.
This module makes use of the below open source projects:
* [aws-managed-grafana](https://github.com/terraform-aws-modules/terraform-aws-managed-service-grafana)
* [aws-managed-prometheus](https://github.com/terraform-aws-modules/terraform-aws-managed-service-prometheus)
See examples using this Terraform modules in the **Amazon ECS** section of [this documentation](https://aws-observability.github.io/terraform-aws-observability-accelerator/)
<!-- BEGINNING OF PRE-COMMIT-TERRAFORM DOCS HOOK -->
## Requirements
| Name | Version |
|------|---------|
| <a name="requirement_terraform"></a> [terraform](#requirement\_terraform) | >= 1.0.0 |
| <a name="requirement_aws"></a> [aws](#requirement\_aws) | >= 5.0.0 |
## Providers
| Name | Version |
|------|---------|
| <a name="provider_aws"></a> [aws](#provider\_aws) | >= 5.0.0 |
## Modules
| Name | Source | Version |
|------|--------|---------|
| <a name="module_managed_grafana_default"></a> [managed\_grafana\_default](#module\_managed\_grafana\_default) | terraform-aws-modules/managed-service-grafana/aws | 2.1.0 |
| <a name="module_managed_prometheus_default"></a> [managed\_prometheus\_default](#module\_managed\_prometheus\_default) | terraform-aws-modules/managed-service-prometheus/aws | 2.2.2 |
## Resources
| Name | Type |
|------|------|
| [aws_ecs_service.adot_ecs_prometheus](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/resources/ecs_service) | resource |
| [aws_ecs_task_definition.adot_ecs_prometheus](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/resources/ecs_task_definition) | resource |
| [aws_ssm_parameter.adot_config](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/resources/ssm_parameter) | resource |
| [aws_region.current](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/data-sources/region) | data source |
## Inputs
| Name | Description | Type | Default | Required |
|------|-------------|------|---------|:--------:|
| <a name="input_aws_ecs_cluster_name"></a> [aws\_ecs\_cluster\_name](#input\_aws\_ecs\_cluster\_name) | Name of your ECS cluster | `string` | n/a | yes |
| <a name="input_container_name"></a> [container\_name](#input\_container\_name) | Container Name for Adot | `string` | `"adot_new"` | no |
| <a name="input_create_managed_grafana_ws"></a> [create\_managed\_grafana\_ws](#input\_create\_managed\_grafana\_ws) | Creates a Workspace for Amazon Managed Grafana | `bool` | `true` | no |
| <a name="input_create_managed_prometheus_ws"></a> [create\_managed\_prometheus\_ws](#input\_create\_managed\_prometheus\_ws) | Creates a Workspace for Amazon Managed Prometheus | `bool` | `true` | no |
| <a name="input_ecs_adot_cpu"></a> [ecs\_adot\_cpu](#input\_ecs\_adot\_cpu) | CPU to be allocated for the ADOT ECS TASK | `string` | `"256"` | no |
| <a name="input_ecs_adot_mem"></a> [ecs\_adot\_mem](#input\_ecs\_adot\_mem) | Memory to be allocated for the ADOT ECS TASK | `string` | `"512"` | no |
| <a name="input_ecs_metrics_collection_interval"></a> [ecs\_metrics\_collection\_interval](#input\_ecs\_metrics\_collection\_interval) | Collection interval for ecs metrics | `string` | `"15s"` | no |
| <a name="input_execution_role_arn"></a> [execution\_role\_arn](#input\_execution\_role\_arn) | ARN of the IAM Execution Role | `string` | n/a | yes |
| <a name="input_otel_image_ver"></a> [otel\_image\_ver](#input\_otel\_image\_ver) | Otel Docker Image version | `string` | `"v0.31.0"` | no |
| <a name="input_otlp_grpc_endpoint"></a> [otlp\_grpc\_endpoint](#input\_otlp\_grpc\_endpoint) | otlpGrpcEndpoint | `string` | `"0.0.0.0:4317"` | no |
| <a name="input_otlp_http_endpoint"></a> [otlp\_http\_endpoint](#input\_otlp\_http\_endpoint) | otlpHttpEndpoint | `string` | `"0.0.0.0:4318"` | no |
| <a name="input_refresh_interval"></a> [refresh\_interval](#input\_refresh\_interval) | Refresh interval for ecs\_observer | `string` | `"60s"` | no |
| <a name="input_task_role_arn"></a> [task\_role\_arn](#input\_task\_role\_arn) | ARN of the IAM Task Role | `string` | n/a | yes |
## Outputs
| Name | Description |
|------|-------------|
| <a name="output_grafana_workspace_endpoint"></a> [grafana\_workspace\_endpoint](#output\_grafana\_workspace\_endpoint) | The endpoint of the Grafana workspace |
| <a name="output_grafana_workspace_id"></a> [grafana\_workspace\_id](#output\_grafana\_workspace\_id) | The ID of the Grafana workspace |
| <a name="output_prometheus_workspace_id"></a> [prometheus\_workspace\_id](#output\_prometheus\_workspace\_id) | Identifier of the workspace |
| <a name="output_prometheus_workspace_prometheus_endpoint"></a> [prometheus\_workspace\_prometheus\_endpoint](#output\_prometheus\_workspace\_prometheus\_endpoint) | Prometheus endpoint available for this workspace |
<!-- END OF PRE-COMMIT-TERRAFORM DOCS HOOK -->
+130
View File
@@ -0,0 +1,130 @@
extensions:
sigv4auth:
region: "${aws_region}"
service: "aps"
ecs_observer: # extension type is ecs_observer
cluster_name: "${cluster_name}" # cluster name need to configured manually
cluster_region: "${cluster_region}" # region can be configured directly or use AWS_REGION env var
result_file: "/etc/ecs_sd_targets.yaml" # the directory for file must already exists
refresh_interval: ${refresh_interval}
job_label_name: prometheus_job
# JMX
docker_labels:
- port_label: "ECS_PROMETHEUS_EXPORTER_PORT"
receivers:
otlp:
protocols:
grpc:
endpoint: ${otlp_grpc_endpoint}
http:
endpoint: ${otlp_http_endpoint}
prometheus:
config:
scrape_configs:
- job_name: "ecssd"
file_sd_configs:
- files:
- "/etc/ecs_sd_targets.yaml"
relabel_configs:
- source_labels: [__meta_ecs_cluster_name]
action: replace
target_label: ClusterName
- source_labels: [__meta_ecs_service_name]
action: replace
target_label: ServiceName
- source_labels: [__meta_ecs_task_definition_family]
action: replace
target_label: TaskDefinitionFamily
- source_labels: [__meta_ecs_task_launch_type]
action: replace
target_label: LaunchType
- source_labels: [__meta_ecs_container_name]
action: replace
target_label: container_name
- action: labelmap
regex: ^__meta_ecs_container_labels_(.+)$
replacement: "$$1"
awsecscontainermetrics:
collection_interval: ${ecs_metrics_collection_interval}
processors:
resource:
attributes:
- key: receiver
value: "prometheus"
action: insert
filter:
metrics:
include:
match_type: strict
metric_names:
- ecs.task.memory.utilized
- ecs.task.memory.reserved
- ecs.task.memory.usage
- ecs.task.cpu.utilized
- ecs.task.cpu.reserved
- ecs.task.cpu.usage.vcpu
- ecs.task.network.rate.rx
- ecs.task.network.rate.tx
- ecs.task.storage.read_bytes
- ecs.task.storage.write_bytes
metricstransform:
transforms:
- include: ".*"
match_type: regexp
action: update
operations:
- label: prometheus_job
new_label: job
action: update_label
- include: ecs.task.memory.utilized
action: update
new_name: MemoryUtilized
- include: ecs.task.memory.reserved
action: update
new_name: MemoryReserved
- include: ecs.task.memory.usage
action: update
new_name: MemoryUsage
- include: ecs.task.cpu.utilized
action: update
new_name: CpuUtilized
- include: ecs.task.cpu.reserved
action: update
new_name: CpuReserved
- include: ecs.task.cpu.usage.vcpu
action: update
new_name: CpuUsage
- include: ecs.task.network.rate.rx
action: update
new_name: NetworkRxBytes
- include: ecs.task.network.rate.tx
action: update
new_name: NetworkTxBytes
- include: ecs.task.storage.read_bytes
action: update
new_name: StorageReadBytes
- include: ecs.task.storage.write_bytes
action: update
new_name: StorageWriteBytes
exporters:
prometheusremotewrite:
endpoint: "${amp_remote_write_ep}"
auth:
authenticator: sigv4auth
logging:
loglevel: debug
service:
extensions: [ecs_observer, sigv4auth]
pipelines:
metrics:
receivers: [prometheus]
processors: [resource, metricstransform]
exporters: [prometheusremotewrite]
metrics/ecs:
receivers: [awsecscontainermetrics]
processors: [filter]
exporters: [logging, prometheusremotewrite]
+32
View File
@@ -0,0 +1,32 @@
data "aws_region" "current" {}
locals {
region = data.aws_region.current.name
name = "amg-ex-${replace(basename(path.cwd), "_", "-")}"
description = "AWS Managed Grafana service for ${local.name}"
prometheus_ws_endpoint = module.managed_prometheus_default[0].workspace_prometheus_endpoint
default_otel_values = {
aws_region = data.aws_region.current.name
cluster_name = var.aws_ecs_cluster_name
cluster_region = data.aws_region.current.name
refresh_interval = var.refresh_interval
ecs_metrics_collection_interval = var.ecs_metrics_collection_interval
amp_remote_write_ep = "${local.prometheus_ws_endpoint}api/v1/remote_write"
otlp_grpc_endpoint = var.otlp_grpc_endpoint
otlp_http_endpoint = var.otlp_http_endpoint
}
ssm_param_value = yamlencode(
templatefile("${path.module}/configs/config.yaml", local.default_otel_values)
)
container_def_default_values = {
container_name = var.container_name
otel_image_ver = var.otel_image_ver
aws_region = data.aws_region.current.name
}
container_definitions = templatefile("${path.module}/task-definitions/otel_collector.json", local.container_def_default_values)
}
+53
View File
@@ -0,0 +1,53 @@
# SSM Parameter for storing and distrivuting the ADOT config
resource "aws_ssm_parameter" "adot_config" {
name = "/terraform-aws-observability/otel_collector_config"
description = "SSM parameter for aws-observability-accelerator/otel-collector-config"
type = "String"
value = local.ssm_param_value
tier = "Intelligent-Tiering"
}
############################################
# Managed Grafana and Prometheus Module
############################################
module "managed_grafana_default" {
count = var.create_managed_grafana_ws ? 1 : 0
source = "terraform-aws-modules/managed-service-grafana/aws"
version = "2.1.0"
name = "${local.name}-default"
associate_license = false
}
module "managed_prometheus_default" {
count = var.create_managed_prometheus_ws ? 1 : 0
source = "terraform-aws-modules/managed-service-prometheus/aws"
version = "2.2.2"
workspace_alias = "${local.name}-default"
}
###########################################
# Task Definition for ADOT ECS Prometheus
###########################################
resource "aws_ecs_task_definition" "adot_ecs_prometheus" {
family = "adot_prometheus_td"
task_role_arn = var.task_role_arn
execution_role_arn = var.execution_role_arn
network_mode = "bridge"
requires_compatibilities = ["EC2"]
cpu = var.ecs_adot_cpu
memory = var.ecs_adot_mem
container_definitions = local.container_definitions
}
############################################
# ECS Service
############################################
resource "aws_ecs_service" "adot_ecs_prometheus" {
name = "adot_prometheus_svc"
cluster = var.aws_ecs_cluster_name
task_definition = aws_ecs_task_definition.adot_ecs_prometheus.arn
desired_count = 1
}
+19
View File
@@ -0,0 +1,19 @@
output "grafana_workspace_id" {
description = "The ID of the Grafana workspace"
value = try(module.managed_grafana_default[0].workspace_id, "")
}
output "grafana_workspace_endpoint" {
description = "The endpoint of the Grafana workspace"
value = try(module.managed_grafana_default[0].workspace_endpoint, "")
}
output "prometheus_workspace_id" {
description = "Identifier of the workspace"
value = try(module.managed_prometheus_default[0].id, "")
}
output "prometheus_workspace_prometheus_endpoint" {
description = "Prometheus endpoint available for this workspace"
value = try(module.managed_prometheus_default[0].prometheus_endpoint, "")
}
@@ -0,0 +1,21 @@
[
{
"name": "${container_name}",
"image": "amazon/aws-otel-collector:${otel_image_ver}",
"secrets": [
{
"name": "AOT_CONFIG_CONTENT",
"valueFrom": "/terraform-aws-observability/otel_collector_config"
}
],
"logConfiguration": {
"logDriver": "awslogs",
"options": {
"awslogs-create-group": "True",
"awslogs-group": "/adot/collector",
"awslogs-region": "${aws_region}",
"awslogs-stream-prefix": "ecs-prometheus"
}
}
}
]
+75
View File
@@ -0,0 +1,75 @@
variable "aws_ecs_cluster_name" {
description = "Name of your ECS cluster"
type = string
}
variable "task_role_arn" {
description = "ARN of the IAM Task Role"
type = string
}
variable "execution_role_arn" {
description = "ARN of the IAM Execution Role"
type = string
}
variable "ecs_adot_cpu" {
description = "CPU to be allocated for the ADOT ECS TASK"
type = string
default = "256"
}
variable "ecs_adot_mem" {
description = "Memory to be allocated for the ADOT ECS TASK"
type = string
default = "512"
}
variable "create_managed_grafana_ws" {
description = "Creates a Workspace for Amazon Managed Grafana"
type = bool
default = true
}
variable "create_managed_prometheus_ws" {
description = "Creates a Workspace for Amazon Managed Prometheus"
type = bool
default = true
}
variable "refresh_interval" {
description = "Refresh interval for ecs_observer"
type = string
default = "60s"
}
variable "ecs_metrics_collection_interval" {
description = "Collection interval for ecs metrics"
type = string
default = "15s"
}
variable "otlp_grpc_endpoint" {
description = "otlpGrpcEndpoint"
type = string
default = "0.0.0.0:4317"
}
variable "otlp_http_endpoint" {
description = "otlpHttpEndpoint"
type = string
default = "0.0.0.0:4318"
}
variable "container_name" {
description = "Container Name for Adot"
type = string
default = "adot_new"
}
variable "otel_image_ver" {
description = "Otel Docker Image version"
type = string
default = "v0.31.0"
}
+10
View File
@@ -0,0 +1,10 @@
terraform {
required_version = ">= 1.0.0"
required_providers {
aws = {
source = "hashicorp/aws"
version = ">= 5.0.0"
}
}
}