EKS Cross Account Observability using central AMP (#213)

* Added example for multi-cluster eks-monitoring and made changes to eks-monitoring module to allow cross-cluster IRSA

* Updated eks-monitoring module's README to add the variable description

* Hard-coded grafana license type in eks-cross-cluster-with-amp/main.tf

* Added README page for eks-cross-account-with-central-amp

* Extracted out the eks/amg creation and modified it to use existing resources

* Update README.md to add multiaccount dashboard png

* Updated README.md and multiaccount.md to add eks multiaccount png

* Updated eks-cross-cluster-with-amp example to disable dashboard creation for cluster 2

* Pre commit changed committed

* Fixed cross-account-observability docs and README.md and added variable for amp_workpace_alias

* Removed extra spacing in multiaccount.md and added precommit suggested changes

* Modified cross-account-observability example to change cross-account-amp-role to snake_case

* Capitalized Terraform string in multiaccount.md and converted iam-role-attach to snake_case

---------

Co-authored-by: Rodrigue Koffi <bonclay7@users.noreply.github.com>
This commit is contained in:
Venkat Penmetsa
2023-09-01 11:55:44 -05:00
committed by GitHub
parent a05e82b220
commit d7daeb8c22
16 changed files with 681 additions and 3 deletions
@@ -0,0 +1,126 @@
# AWS EKS Cross Account Observability
This example shows how to use the [AWS Observability Accelerator](https://github.com/aws-observability/terraform-aws-observability-accelerator), with two or more EKS cluster in multiple AWS accounts and verify the collected metrics from all the clusters in the dashboards of a common `Amazon Managed Grafana` workspace in a central monitoring account.
## Prerequisites
#### 1. Cross Account IAM access
In order to create/modify resources across multiple AWS accounts, this Terraform example implements the cross-account IAM role assumption. You will need separate IAM roles in all 3 AWS accounts, and each of these IAM roles should have the below specified trust-relationship so that your local AWS user/role will be able to assume them during the terraform execution.
```
{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Principal": {
"AWS": "<local-aws-user/role-arn>"
},
"Action": "sts:AssumeRole",
"Condition": {}
}
]
}
```
> [!NOTE]
> The IAM roles in Account 1 and Account 2 (EKS cluster accounts) should have permissions to perform kubernetes API operations against your EKS clusters. For more info, please review documentation for [enabling IAM principal access to your clusters](https://docs.aws.amazon.com/eks/latest/userguide/add-user-role.html)
#### 2. EKS clusters in multiple AWS Accounts
Using the example [eks-cluster-with-vpc](../../examples/eks-cluster-with-vpc/), create two EKS clusters with the below names in two different AWS accounts:
1. `eks-cluster-1` (Account 1)
2. `eks-cluster-2` (Account 2)
Update the cluster names and their corresponding region names in the `variables.tf` file along with the corresponding IAM role ARNs that can be assumed by terraform to perform cross-account API operations.
#### 3. Amazon Managed Grafana (AMG) workspace
To run this example you need an existing Amazon Managed Grafana (AMG) workspace. If not, you can create a new AMG workspace by following the [Getting Started with Amazon Managed Grafana](https://docs.aws.amazon.com/grafana/latest/userguide/getting-started-with-AMG.html) documentation.
Add the Grafana Workspace ID and its corresponding region name in the `variables.tf` file along with the corresponding IAM role ARN that can be assumed by terraform to perform cross-account API operations.
!!! note
You can obtain the AMG Workspace ID based on its URL. For the URL `https://g-xyz.grafana-workspace.eu-central-1.amazonaws.com`, the workspace ID would be `g-xyz`
## Setup
#### 1. Download sources and initialize Terraform
```sh
git clone https://github.com/aws-observability/terraform-aws-observability-accelerator.git
cd terraform-aws-observability-accelerator/examples/eks-cross-account-with-central-amp
terraform init
```
#### 2. Deploy
By looking at the `variables.tf`, you will notice there are two EKS clusters targeted for deployment by the names/ids:
1. `eks-cluster-1`
2. `eks-cluster-2`
While installing the observability settings for the EKS cluster specified in variable `cluster_one.name`, Terraform also sets up:
* Creates an `Amazon Managed Prometheus Workspace`
* Dashboard folder and files in provided `Amazon Managed Grafana Workspace`
!!! warning
To override the defaults, create a `terraform.tfvars` and change the default values of the variables.
Run the following command to deploy
```sh
terraform apply --auto-approve
```
## Verifying Multi Account Observability
One you have successfully run the above setup, you should be able to see dashboards similar to the images shown below in `Amazon Managed Grafana` workspace.
You will notice that you are able to use the `cluster` dropdown to filter the dashboards to metrics collected from a specific EKS cluster.
![eks-cross-account-1](https://github.com/veekaly/terraform-aws-observability-accelerator/assets/119073483/96a68eb1-4fb7-4a6b-bd4a-15f4f6ac7565)
![eks-cross-account-2](https://github.com/veekaly/terraform-aws-observability-accelerator/assets/119073483/1373b834-1082-4a63-98b9-2b90fb32eada)
## Cleanup
To clean up entirely, run the following command:
```sh
terraform destroy --auto-approve
```
@@ -0,0 +1,19 @@
data "aws_eks_cluster_auth" "eks_one" {
name = var.cluster_one.name
provider = aws.eks_cluster_one
}
data "aws_eks_cluster_auth" "eks_two" {
name = var.cluster_two.name
provider = aws.eks_cluster_two
}
data "aws_eks_cluster" "eks_one" {
name = var.cluster_one.name
provider = aws.eks_cluster_one
}
data "aws_eks_cluster" "eks_two" {
name = var.cluster_two.name
provider = aws.eks_cluster_two
}
@@ -0,0 +1,73 @@
data "aws_caller_identity" "monitoring" {
provider = aws.central_monitoring
}
resource "aws_iam_policy" "irsa_assume_role_policy_one" {
provider = aws.eks_cluster_one
name = "${var.cluster_one.name}-irsa_assume_role_policy"
path = "/"
description = "This role allows the IRSA role to assume the cross-account role for AMP access"
policy = jsonencode({
Version = "2012-10-17"
Statement = [
{
Action = [
"sts:AssumeRole",
]
Effect = "Allow"
Resource = "arn:aws:iam::${data.aws_caller_identity.monitoring.account_id}:role/${local.amp_workspace_alias}-role-for-cross-account"
},
]
})
}
resource "aws_iam_policy" "irsa_assume_role_policy_two" {
provider = aws.eks_cluster_two
name = "${var.cluster_two.name}-irsa_assume_role_policy"
path = "/"
description = "This role allows the IRSA role to assume the cross-account role for AMP access"
policy = jsonencode({
Version = "2012-10-17"
Statement = [
{
Action = [
"sts:AssumeRole",
]
Effect = "Allow"
Resource = "arn:aws:iam::${data.aws_caller_identity.monitoring.account_id}:role/${local.amp_workspace_alias}-role-for-cross-account"
},
]
})
}
resource "aws_iam_role" "cross_account_amp_role" {
provider = aws.central_monitoring
name = "${local.amp_workspace_alias}-role-for-cross-account"
assume_role_policy = <<EOF
{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Principal": {
"AWS": [
"${module.eks_monitoring_one.adot_irsa_arn}",
"${module.eks_monitoring_two.adot_irsa_arn}"
]
},
"Action": "sts:AssumeRole",
"Condition": {}
}
]
}
EOF
}
resource "aws_iam_role_policy_attachment" "role_attach" {
provider = aws.central_monitoring
role = aws_iam_role.cross_account_amp_role.name
policy_arn = "arn:aws:iam::aws:policy/AmazonPrometheusRemoteWriteAccess"
}
@@ -0,0 +1,141 @@
locals {
amp_workspace_alias = var.monitoring.amp_workspace_alias
}
###########################################################################
# EKS Monitoring Addon for cluster one #
###########################################################################
module "eks_monitoring_one" {
source = "../../modules/eks-monitoring"
# source = "github.com/aws-observability/terraform-aws-observability-accelerator//modules/eks-monitoring?ref=v2.0.0"
providers = {
aws = aws.eks_cluster_one
helm = helm.eks_cluster_one
kubernetes = kubernetes.eks_cluster_one
kubectl = kubectl.eks_cluster_one
}
eks_cluster_id = var.cluster_one.name
enable_amazon_eks_adot = true
enable_cert_manager = true
enable_fluxcd = true
enable_external_secrets = true
enable_dashboards = true
enable_java = true
enable_nginx = true
enable_node_exporter = true
# Set to false for cross-cluster observability
enable_alerting_rules = false
enable_recording_rules = false
grafana_api_key = aws_grafana_workspace_api_key.key.key
target_secret_name = "grafana-admin-credentials"
target_secret_namespace = "grafana-operator"
grafana_url = module.aws_observability_accelerator.managed_grafana_workspace_endpoint
managed_prometheus_workspace_id = module.aws_observability_accelerator.managed_prometheus_workspace_id
managed_prometheus_workspace_endpoint = module.aws_observability_accelerator.managed_prometheus_workspace_endpoint
managed_prometheus_workspace_region = module.aws_observability_accelerator.managed_prometheus_workspace_region
managed_prometheus_cross_account_role = aws_iam_role.cross_account_amp_role.arn
irsa_iam_additional_policies = [aws_iam_policy.irsa_assume_role_policy_one.arn]
# optional, defaults to 60s interval and 15s timeout
prometheus_config = {
global_scrape_interval = "60s"
global_scrape_timeout = "15s"
}
enable_logs = true
depends_on = [
module.aws_observability_accelerator
]
}
###########################################################################
# EKS Monitoring Addon for cluster two #
###########################################################################
module "eks_monitoring_two" {
source = "../../modules/eks-monitoring"
# source = "github.com/aws-observability/terraform-aws-observability-accelerator//modules/eks-monitoring?ref=v2.0.0"
providers = {
aws = aws.eks_cluster_two
helm = helm.eks_cluster_two
kubernetes = kubernetes.eks_cluster_two
kubectl = kubectl.eks_cluster_two
}
eks_cluster_id = var.cluster_two.name
enable_amazon_eks_adot = true
enable_cert_manager = true
enable_fluxcd = false
enable_external_secrets = false
enable_dashboards = false
enable_node_exporter = true
# Set to false for cross-cluster observability
enable_alerting_rules = false
enable_recording_rules = false
grafana_api_key = aws_grafana_workspace_api_key.key.key
target_secret_name = "grafana-admin-credentials"
target_secret_namespace = "grafana-operator"
grafana_url = module.aws_observability_accelerator.managed_grafana_workspace_endpoint
managed_prometheus_workspace_id = module.aws_observability_accelerator.managed_prometheus_workspace_id
managed_prometheus_workspace_endpoint = module.aws_observability_accelerator.managed_prometheus_workspace_endpoint
managed_prometheus_workspace_region = module.aws_observability_accelerator.managed_prometheus_workspace_region
managed_prometheus_cross_account_role = aws_iam_role.cross_account_amp_role.arn
irsa_iam_additional_policies = [aws_iam_policy.irsa_assume_role_policy_two.arn]
# optional, defaults to 60s interval and 15s timeout
prometheus_config = {
global_scrape_interval = "60s"
global_scrape_timeout = "15s"
}
enable_logs = true
depends_on = [
module.aws_observability_accelerator
]
}
###########################################################################
# AMP and Grafana resources #
###########################################################################
resource "aws_grafana_workspace_api_key" "key" {
provider = aws.central_monitoring
key_name = "terraform-api-key"
key_role = "ADMIN"
seconds_to_live = 86400
workspace_id = var.monitoring.managed_grafana_id
}
module "managed_service_prometheus" {
source = "terraform-aws-modules/managed-service-prometheus/aws"
version = "2.2.2"
providers = {
aws = aws.central_monitoring
}
workspace_alias = local.amp_workspace_alias
}
module "aws_observability_accelerator" {
source = "../../../terraform-aws-observability-accelerator"
aws_region = var.monitoring.region
enable_managed_prometheus = false
enable_alertmanager = false
managed_prometheus_workspace_region = var.monitoring.region
managed_prometheus_workspace_id = module.managed_service_prometheus.workspace_id
managed_grafana_workspace_id = var.monitoring.managed_grafana_id
providers = {
aws = aws.central_monitoring
}
}
@@ -0,0 +1,4 @@
output "amp_workspace_id" {
description = "Identifier of the AMP workspace"
value = module.managed_service_prometheus.workspace_id
}
@@ -0,0 +1,95 @@
###### AWS Providers ######
provider "aws" {
region = var.cluster_one.region
alias = "eks_cluster_one"
assume_role {
role_arn = var.cluster_one.tf_role
}
}
provider "aws" {
region = var.cluster_two.region
alias = "eks_cluster_two"
assume_role {
role_arn = var.cluster_two.tf_role
}
}
provider "aws" {
region = var.monitoring.region
alias = "central_monitoring"
assume_role {
role_arn = var.monitoring.tf_role
}
}
###### Helm Providers ######
provider "helm" {
alias = "eks_cluster_one"
kubernetes {
host = data.aws_eks_cluster.eks_one.endpoint
cluster_ca_certificate = base64decode(data.aws_eks_cluster.eks_one.certificate_authority[0].data)
exec {
api_version = "client.authentication.k8s.io/v1beta1"
args = ["eks", "get-token", "--role-arn", var.cluster_one.tf_role, "--cluster-name", var.cluster_one.name]
command = "aws"
}
}
}
provider "helm" {
alias = "eks_cluster_two"
kubernetes {
host = data.aws_eks_cluster.eks_two.endpoint
cluster_ca_certificate = base64decode(data.aws_eks_cluster.eks_two.certificate_authority[0].data)
exec {
api_version = "client.authentication.k8s.io/v1beta1"
args = ["eks", "get-token", "--role-arn", var.cluster_two.tf_role, "--cluster-name", var.cluster_two.name]
command = "aws"
}
}
}
###### Kubernetes Providers ######
provider "kubernetes" {
alias = "eks_cluster_one"
host = data.aws_eks_cluster.eks_one.endpoint
cluster_ca_certificate = base64decode(data.aws_eks_cluster.eks_one.certificate_authority[0].data)
exec {
api_version = "client.authentication.k8s.io/v1beta1"
args = ["eks", "get-token", "--role-arn", var.cluster_one.tf_role, "--cluster-name", var.cluster_one.name]
command = "aws"
}
}
provider "kubernetes" {
alias = "eks_cluster_two"
host = data.aws_eks_cluster.eks_two.endpoint
cluster_ca_certificate = base64decode(data.aws_eks_cluster.eks_two.certificate_authority[0].data)
exec {
api_version = "client.authentication.k8s.io/v1beta1"
args = ["eks", "get-token", "--role-arn", var.cluster_two.tf_role, "--cluster-name", var.cluster_two.name]
command = "aws"
}
}
provider "kubectl" {
alias = "eks_cluster_one"
apply_retry_count = 30
host = data.aws_eks_cluster.eks_one.endpoint
cluster_ca_certificate = base64decode(data.aws_eks_cluster.eks_one.certificate_authority[0].data)
load_config_file = false
token = data.aws_eks_cluster_auth.eks_one.token
}
provider "kubectl" {
alias = "eks_cluster_two"
apply_retry_count = 30
host = data.aws_eks_cluster.eks_two.endpoint
cluster_ca_certificate = base64decode(data.aws_eks_cluster.eks_two.certificate_authority[0].data)
load_config_file = false
token = data.aws_eks_cluster_auth.eks_two.token
}
@@ -0,0 +1,43 @@
variable "cluster_one" {
description = "Input for your first EKS Cluster"
type = object({
name = string
region = string
tf_role = string
})
default = {
name = "eks-cluster-1"
region = "us-east-1"
tf_role = "<iam-role-in-eks-cluster-1-account>"
}
}
variable "cluster_two" {
description = "Input for your second EKS Cluster"
type = object({
name = string
region = string
tf_role = string
})
default = {
name = "eks-cluster-2"
region = "us-east-1"
tf_role = "<iam-role-in-eks-cluster-2-account>"
}
}
variable "monitoring" {
description = "Input for your AMP and AMG workspaces"
type = object({
managed_grafana_id = string
amp_workspace_alias = string
region = string
tf_role = string
})
default = {
managed_grafana_id = "<grafana-ws-id>"
amp_workspace_alias = "aws-observability-accelerator"
region = "<grafana-ws-region>"
tf_role = "<iam-role-in-grafana-ws-account>"
}
}
@@ -0,0 +1,25 @@
terraform {
required_version = ">= 1.3.9"
required_providers {
aws = {
source = "hashicorp/aws"
version = ">= 4.55.0"
configuration_aliases = [aws.eks_cluster_one, aws.eks_cluster_two, aws.central_monitoring]
}
kubernetes = {
source = "hashicorp/kubernetes"
version = ">= 2.18.0"
configuration_aliases = [kubernetes.eks_cluster_one, kubernetes.eks_cluster_two]
}
helm = {
source = "hashicorp/helm"
version = ">= 2.9.0"
configuration_aliases = [helm.eks_cluster_one, helm.eks_cluster_two]
}
kubectl = {
source = "gavinbunney/kubectl"
version = ">= 1.14"
}
}
}