diff --git a/docs/eks/istio.md b/docs/eks/istio.md new file mode 100644 index 0000000..9829e8c --- /dev/null +++ b/docs/eks/istio.md @@ -0,0 +1,174 @@ +# Monitor Istio running on Amazon EKS + +This example demonstrates how to use Terraform modules for AWS Observability Accelerator, EKS Blueprints with the Tetrate Istio Add-on and EKS monitoring for Istio. + +The current example deploys the [AWS Distro for OpenTelemetry Operator](https://docs.aws.amazon.com/eks/latest/userguide/opentelemetry.html) +for Amazon EKS with its requirements and make use of an existing Amazon Managed Grafana workspace. +It creates a new Amazon Managed Service for Prometheus workspace unless provided with an existing one to reuse. + +It uses the `EKS monitoring` [module](../../modules/eks-monitoring/) +to provide an existing EKS cluster with an OpenTelemetry collector, +curated Grafana dashboards, Prometheus alerting and recording rules with multiple +configuration options for Istio. + +## Prerequisites + +Ensure that you have the following tools installed locally: + +1. [aws cli](https://docs.aws.amazon.com/cli/latest/userguide/getting-started-install.html) +2. [kubectl](https://kubernetes.io/docs/tasks/tools/) +3. [terraform](https://learn.hashicorp.com/tutorials/terraform/install-cli) +4. [istioctl](https://istio.io/latest/docs/setup/getting-started/#download) + +## Setup + +This example uses a local terraform state. If you need states to be saved remotely, +on Amazon S3 for example, visit the [terraform remote states](https://www.terraform.io/language/state/remote) documentation + +### 1. Clone the repo using the command below + +``` +git clone https://github.com/aws-observability/terraform-aws-observability-accelerator.git +``` + +### 2. Initialize terraform + +```console +cd examples/eks-istio +terraform init +``` + +### 3. Amazon EKS Cluster + +To run this example, you need to provide your EKS cluster name. +If you don't have a cluster ready, visit [this example](https://aws-observability.github.io/terraform-aws-observability-accelerator/helpers/new-eks-cluster/) +first to create a new one. + +Add your cluster name for `eks_cluster_id="..."` to the `terraform.tfvars` or use an environment variable `export TF_VAR_eks_cluster_id=xxx`. + +### 4. Amazon Managed Grafana workspace + +To run this example you need an Amazon Managed Grafana workspace. If you have +an existing workspace, create an environment variable +`export TF_VAR_managed_grafana_workspace_id=g-xxx`. + +To create a new one, visit [this example](https://aws-observability.github.io/terraform-aws-observability-accelerator/helpers/managed-grafana/). + +> In the URL `https://g-xyz.grafana-workspace.eu-central-1.amazonaws.com`, the workspace ID would be `g-xyz` + +### 5. Grafana API Key + +Amazon Managed Service for Grafana provides a control plane API for generating Grafana API keys. We will provide to Terraform +a short lived API key to run the `apply` or `destroy` command. +Ensure you have necessary IAM permissions (`CreateWorkspaceApiKey, DeleteWorkspaceApiKey`) + +```sh +export TF_VAR_grafana_api_key=`aws grafana create-workspace-api-key --key-name "observability-accelerator-$(date +%s)" --key-role ADMIN --seconds-to-live 1200 --workspace-id $TF_VAR_managed_grafana_workspace_id --query key --output text` +``` + +## Deploy + +Simply run this command to deploy (if using a variable definition file) + +```sh +terraform apply -var-file=terraform.tfvars +``` + +or if you had setup environment variables, run + +```sh +terraform apply +``` + +## Additional configuration + +For the purpose of the example, we have provided default values for some of the variables. + +1. AWS Region + +Specify the AWS Region where the resources will be deployed. Edit the `terraform.tfvars` file and modify `aws_region="..."`. You can also use environement variables `export TF_VAR_aws_region=xxx`. + + +2. Amazon Managed Service for Prometheus workspace + +If you have an existing workspace, add `managed_prometheus_workspace_id=ws-xxx` +or use an environment variable `export TF_VAR_managed_prometheus_workspace_id=ws-xxx`. + +## Visualization + +### 1. Grafana dashboards + +Go to the Dashboards panel of your Grafana workspace. You will see a list of Istio dashboards under the `Observability Accelerator Dashboards` + +image + +Open one of the Istio dasbhoards and you will be able to view its visualization + +image + +### 2. Amazon Managed Service for Prometheus rules and alerts + +Open the Amazon Managed Service for Prometheus console and view the details of your workspace. Under the `Rules management` tab, you will find new rules deployed. + +image + +!!! note + To setup your alert receiver, with Amazon SNS, follow [this documentation](https://docs.aws.amazon.com/prometheus/latest/userguide/AMP-alertmanager-receiver.html) + +## Deploy an example application to visualize metrics + +In this section we will deploy Istio's Bookinfo sample application and extract metrics using the AWS OpenTelemetry collector. When downloading and configuring `istioctl`, there are samples included in the Istio package directory. The deployment files for Bookinfo are found in the `samples` folder. Additional details can be found on Istio's [Getting Started](https://istio.io/latest/docs/setup/getting-started/) documentation + +### 1. Deploy the Bookinfo Application + +1. Using the AWS CLI, configure kubectl so you can connect to your EKS cluster. Update for your region and EKS cluster name +```sh +aws eks update-kubeconfig --region --name +``` +2. Label the default namespace for automatic Istio sidecar injection +```sh +kubectl label namespace default istio-injection=enabled +``` +3. Navigate to the Istio folder location. For example, if using Istio v1.18.2 in Downloads folder: +```sh +cd ~/Downloads/istio-1.18.2 +``` +4. Deploy the Bookinfo sample application +```sh +kubectl apply -f samples/bookinfo/platform/kube/bookinfo.yaml +``` +5. Connect the Bookinfo application with the Istio gateway +```sh +kubectl apply -f samples/bookinfo/networking/bookinfo-gateway.yaml +``` +6. Validate that there are no issues with the Istio configuration +```sh +istioctl analyze +``` +7. Get the DNS name of the load balancer for the Istio gateway +```sh +GATEWAY_URL=$(kubectl get svc istio-ingressgateway -n istio-system -o=jsonpath='{.status.loadBalancer.ingress[0].hostname}') +``` + +### 2. Generate traffic for the Istio Bookinfo sample application + +For the Bookinfo sample application, visit `http://$GATEWAY_URL/productpage` in your web browser. To see trace data, you must send requests to your service. The number of requests depends on Istio’s sampling rate and can be configured using the Telemetry API. With the default sampling rate of 1%, you need to send at least 100 requests before the first trace is visible. To send a 100 requests to the productpage service, use the following command: +```sh +for i in $(seq 1 100); do curl -s -o /dev/null "http://$GATEWAY_URL/productpage"; done +``` + +### 3. Explore the Istio dashboards + +Log back into your Amazon Managed Grafana workspace and navigate to the dashboard side panel. Click on the `Observability Accelerator Dashboards` folder and open the `Istio Service` Dashboard. Use the Service dropdown menu to select the `reviews.default.svc.cluster.local` service. This gives details about metrics for the service, client workloads (workloads that are calling this service), and service workloads (workloads that are providing this service). + +Explore the Istio Control Plane, Mesh, and Performance dashboards as well. + +## Destroy + +To teardown and remove the resources created in this example: + +```sh +kubectl delete -f samples/bookinfo/networking/bookinfo-gateway.yaml +kubectl delete -f samples/bookinfo/platform/kube/bookinfo.yaml +terraform destroy +``` diff --git a/examples/eks-istio/README.md b/examples/eks-istio/README.md new file mode 100644 index 0000000..7271915 --- /dev/null +++ b/examples/eks-istio/README.md @@ -0,0 +1,57 @@ +# Existing Cluster with the AWS Observability accelerator base module, Tetrate Istio Add-on and Istio monitoring + +View the full documentation for this example [here](https://aws-observability.github.io/terraform-aws-observability-accelerator/eks/istio) + + +## Requirements + +| Name | Version | +|------|---------| +| [terraform](#requirement\_terraform) | >= 1.1.0 | +| [aws](#requirement\_aws) | >= 4.0.0 | +| [helm](#requirement\_helm) | >= 2.4.1 | +| [kubectl](#requirement\_kubectl) | >= 1.14 | +| [kubernetes](#requirement\_kubernetes) | >= 2.10 | + +## Providers + +| Name | Version | +|------|---------| +| [aws](#provider\_aws) | >= 4.0.0 | + +## Modules + +| Name | Source | Version | +|------|--------|---------| +| [aws\_observability\_accelerator](#module\_aws\_observability\_accelerator) | ../../ | n/a | +| [eks\_blueprints\_kubernetes\_addons](#module\_eks\_blueprints\_kubernetes\_addons) | github.com/aws-ia/terraform-aws-eks-blueprints//modules/kubernetes-addons | v4.32.0 | +| [eks\_monitoring](#module\_eks\_monitoring) | ../../modules/eks-monitoring | n/a | + +## Resources + +| Name | Type | +|------|------| +| [aws_eks_cluster.this](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/data-sources/eks_cluster) | data source | +| [aws_eks_cluster_auth.this](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/data-sources/eks_cluster_auth) | data source | + +## Inputs + +| Name | Description | Type | Default | Required | +|------|-------------|------|---------|:--------:| +| [aws\_region](#input\_aws\_region) | AWS Region | `string` | n/a | yes | +| [eks\_cluster\_id](#input\_eks\_cluster\_id) | Name of the EKS cluster | `string` | `"eks-cluster-with-vpc"` | no | +| [enable\_dashboards](#input\_enable\_dashboards) | Enables or disables curated dashboards. Dashboards are managed by the Grafana Operator | `bool` | `true` | no | +| [grafana\_api\_key](#input\_grafana\_api\_key) | API key for authorizing the Grafana provider to make changes to Amazon Managed Grafana | `string` | n/a | yes | +| [managed\_grafana\_workspace\_id](#input\_managed\_grafana\_workspace\_id) | Amazon Managed Grafana Workspace ID | `string` | n/a | yes | +| [managed\_prometheus\_workspace\_id](#input\_managed\_prometheus\_workspace\_id) | Amazon Managed Service for Prometheus Workspace ID | `string` | `""` | no | + +## Outputs + +| Name | Description | +|------|-------------| +| [aws\_region](#output\_aws\_region) | AWS Region | +| [eks\_cluster\_id](#output\_eks\_cluster\_id) | EKS Cluster Id | +| [eks\_cluster\_version](#output\_eks\_cluster\_version) | EKS Cluster version | +| [managed\_prometheus\_workspace\_endpoint](#output\_managed\_prometheus\_workspace\_endpoint) | Amazon Managed Prometheus workspace endpoint | +| [managed\_prometheus\_workspace\_id](#output\_managed\_prometheus\_workspace\_id) | Amazon Managed Prometheus workspace ID | + diff --git a/examples/eks-istio/main.tf b/examples/eks-istio/main.tf new file mode 100644 index 0000000..667df36 --- /dev/null +++ b/examples/eks-istio/main.tf @@ -0,0 +1,121 @@ +provider "aws" { + region = local.region +} + +data "aws_eks_cluster_auth" "this" { + name = var.eks_cluster_id +} + +data "aws_eks_cluster" "this" { + name = var.eks_cluster_id +} + +provider "kubernetes" { + host = local.eks_cluster_endpoint + cluster_ca_certificate = base64decode(data.aws_eks_cluster.this.certificate_authority[0].data) + token = data.aws_eks_cluster_auth.this.token +} + +provider "helm" { + kubernetes { + host = local.eks_cluster_endpoint + cluster_ca_certificate = base64decode(data.aws_eks_cluster.this.certificate_authority[0].data) + token = data.aws_eks_cluster_auth.this.token + } +} + +locals { + region = var.aws_region + eks_cluster_endpoint = data.aws_eks_cluster.this.endpoint + create_new_workspace = var.managed_prometheus_workspace_id == "" ? true : false + tags = { + Source = "github.com/aws-observability/terraform-aws-observability-accelerator" + } +} + +# deploys the base module +module "aws_observability_accelerator" { + source = "../../" + # source = "github.com/aws-observability/terraform-aws-observability-accelerator?ref=v2.0.0" + + aws_region = var.aws_region + + # creates a new Amazon Managed Prometheus workspace, defaults to true + enable_managed_prometheus = local.create_new_workspace + + # reusing existing Amazon Managed Prometheus if specified + managed_prometheus_workspace_id = var.managed_prometheus_workspace_id + + # sets up the Amazon Managed Prometheus alert manager at the workspace level + enable_alertmanager = true + + # reusing existing Amazon Managed Grafana workspace + managed_grafana_workspace_id = var.managed_grafana_workspace_id + + tags = local.tags +} + +module "eks_blueprints_kubernetes_addons" { + source = "github.com/aws-ia/terraform-aws-eks-blueprints//modules/kubernetes-addons?ref=v4.32.0" + + eks_cluster_id = var.eks_cluster_id + #eks_cluster_endpoint = module.eks_blueprints.eks_cluster_endpoint + #eks_oidc_provider = module.eks_blueprints.oidc_provider + #eks_cluster_version = module.eks_blueprints.eks_cluster_version + + # EKS Managed Add-ons + #enable_amazon_eks_vpc_cni = true + #enable_amazon_eks_coredns = true + #enable_amazon_eks_kube_proxy = true + + # Add-ons + enable_metrics_server = true + enable_cluster_autoscaler = true + + # Tetrate Istio Add-on + enable_tetrate_istio = true + + tags = local.tags +} + +module "eks_monitoring" { + source = "../../modules/eks-monitoring" + # source = "github.com/aws-observability/terraform-aws-observability-accelerator//modules/eks-monitoring?ref=v2.0.0" + enable_istio = true + eks_cluster_id = var.eks_cluster_id + + # deploys AWS Distro for OpenTelemetry operator into the cluster + enable_amazon_eks_adot = true + + # reusing existing certificate manager? defaults to true + enable_cert_manager = true + + # deploys external-secrets in to the cluster + enable_external_secrets = true + grafana_api_key = var.grafana_api_key + target_secret_name = "grafana-admin-credentials" + target_secret_namespace = "grafana-operator" + grafana_url = module.aws_observability_accelerator.managed_grafana_workspace_endpoint + + # control the publishing of dashboards by specifying the boolean value for the variable 'enable_dashboards', default is 'true' + enable_dashboards = var.enable_dashboards + + managed_prometheus_workspace_id = module.aws_observability_accelerator.managed_prometheus_workspace_id + + managed_prometheus_workspace_endpoint = module.aws_observability_accelerator.managed_prometheus_workspace_endpoint + managed_prometheus_workspace_region = module.aws_observability_accelerator.managed_prometheus_workspace_region + + # optional, defaults to 60s interval and 15s timeout + prometheus_config = { + global_scrape_interval = "60s" + global_scrape_timeout = "15s" + } + + enable_logs = true + + tags = local.tags + + depends_on = [ + module.aws_observability_accelerator + ] +} diff --git a/examples/eks-istio/outputs.tf b/examples/eks-istio/outputs.tf new file mode 100644 index 0000000..ad1c340 --- /dev/null +++ b/examples/eks-istio/outputs.tf @@ -0,0 +1,24 @@ +output "aws_region" { + description = "AWS Region" + value = module.aws_observability_accelerator.aws_region +} + +output "managed_prometheus_workspace_endpoint" { + description = "Amazon Managed Prometheus workspace endpoint" + value = module.aws_observability_accelerator.managed_prometheus_workspace_endpoint +} + +output "managed_prometheus_workspace_id" { + description = "Amazon Managed Prometheus workspace ID" + value = module.aws_observability_accelerator.managed_prometheus_workspace_id +} + +output "eks_cluster_version" { + description = "EKS Cluster version" + value = module.eks_monitoring.eks_cluster_version +} + +output "eks_cluster_id" { + description = "EKS Cluster Id" + value = module.eks_monitoring.eks_cluster_id +} diff --git a/examples/eks-istio/variables.tf b/examples/eks-istio/variables.tf new file mode 100644 index 0000000..6df0e2e --- /dev/null +++ b/examples/eks-istio/variables.tf @@ -0,0 +1,33 @@ +variable "eks_cluster_id" { + description = "Name of the EKS cluster" + type = string + default = "eks-cluster-with-vpc" +} + +variable "aws_region" { + description = "AWS Region" + type = string +} + +variable "managed_prometheus_workspace_id" { + description = "Amazon Managed Service for Prometheus Workspace ID" + type = string + default = "" +} + +variable "managed_grafana_workspace_id" { + description = "Amazon Managed Grafana Workspace ID" + type = string +} + +variable "grafana_api_key" { + description = "API key for authorizing the Grafana provider to make changes to Amazon Managed Grafana" + type = string + sensitive = true +} + +variable "enable_dashboards" { + description = "Enables or disables curated dashboards. Dashboards are managed by the Grafana Operator" + type = bool + default = true +} diff --git a/examples/eks-istio/versions.tf b/examples/eks-istio/versions.tf new file mode 100644 index 0000000..046beb3 --- /dev/null +++ b/examples/eks-istio/versions.tf @@ -0,0 +1,30 @@ +terraform { + required_version = ">= 1.1.0" + + required_providers { + aws = { + source = "hashicorp/aws" + version = ">= 4.0.0" + } + kubernetes = { + source = "hashicorp/kubernetes" + version = ">= 2.10" + } + kubectl = { + source = "gavinbunney/kubectl" + version = ">= 1.14" + } + helm = { + source = "hashicorp/helm" + version = ">= 2.4.1" + } + } + + # ## Used for end-to-end testing on project; update to suit your needs + # backend "s3" { + # bucket = "aws-observability-accelerator-terraform-states" + # region = "us-west-2" + # key = "e2e/eks-istio/terraform.tfstate" + # } + +} diff --git a/mkdocs.yml b/mkdocs.yml index d7245d9..7dffeec 100644 --- a/mkdocs.yml +++ b/mkdocs.yml @@ -30,6 +30,7 @@ nav: - Multicluster monitoring: eks/multicluster.md - Java/JMX: eks/java.md - Nginx: eks/nginx.md + - Istio: eks/istio.md - Viewing logs: eks/logs.md - Teardown: eks/destroy.md - Monitoring Managed Service for Prometheus Workspaces: workloads/managed-prometheus.md diff --git a/modules/eks-monitoring/README.md b/modules/eks-monitoring/README.md index a3c4f17..0269284 100644 --- a/modules/eks-monitoring/README.md +++ b/modules/eks-monitoring/README.md @@ -40,6 +40,7 @@ See examples using this Terraform modules in the **Amazon EKS** section of [this | [external\_secrets](#module\_external\_secrets) | ./add-ons/external-secrets | n/a | | [fluentbit\_logs](#module\_fluentbit\_logs) | ./add-ons/aws-for-fluentbit | n/a | | [helm\_addon](#module\_helm\_addon) | github.com/aws-ia/terraform-aws-eks-blueprints//modules/kubernetes-addons/helm-addon | v4.32.0 | +| [istio\_monitoring](#module\_istio\_monitoring) | ./patterns/istio | n/a | | [java\_monitoring](#module\_java\_monitoring) | ./patterns/java | n/a | | [nginx\_monitoring](#module\_nginx\_monitoring) | ./patterns/nginx | n/a | | [operator](#module\_operator) | ./add-ons/adot-operator | n/a | @@ -76,6 +77,7 @@ See examples using this Terraform modules in the **Amazon EKS** section of [this | [enable\_external\_secrets](#input\_enable\_external\_secrets) | Installs External Secrets to EKS Cluster | `bool` | `true` | no | | [enable\_fluxcd](#input\_enable\_fluxcd) | Enables or disables FluxCD. Disabling this might affect some data in the dashboards | `bool` | `true` | no | | [enable\_grafana\_operator](#input\_enable\_grafana\_operator) | Deploys Grafana Operator to EKS Cluster | `bool` | `true` | no | +| [enable\_istio](#input\_enable\_istio) | Enable ISTIO workloads monitoring, alerting and default dashboards | `bool` | `false` | no | | [enable\_java](#input\_enable\_java) | Enable Java workloads monitoring, alerting and default dashboards | `bool` | `false` | no | | [enable\_kube\_state\_metrics](#input\_enable\_kube\_state\_metrics) | Enables or disables Kube State metrics exporter. Disabling this might affect some data in the dashboards | `bool` | `true` | no | | [enable\_logs](#input\_enable\_logs) | Using AWS For FluentBit to collect cluster and application logs to Amazon CloudWatch | `bool` | `true` | no | @@ -101,6 +103,7 @@ See examples using this Terraform modules in the **Amazon EKS** section of [this | [helm\_config](#input\_helm\_config) | Helm Config for Prometheus | `any` | `{}` | no | | [irsa\_iam\_permissions\_boundary](#input\_irsa\_iam\_permissions\_boundary) | IAM permissions boundary for IRSA roles | `string` | `null` | no | | [irsa\_iam\_role\_path](#input\_irsa\_iam\_role\_path) | IAM role path for IRSA roles | `string` | `"/"` | no | +| [istio\_config](#input\_istio\_config) | Configuration object for ISTIO monitoring |
object({
enable_alerting_rules = bool
enable_recording_rules = bool
enable_dashboards = bool
scrape_sample_limit = number

flux_gitrepository_name = string
flux_gitrepository_url = string
flux_gitrepository_branch = string
flux_kustomization_name = string
flux_kustomization_path = string

grafana_url = string
grafana_istio_cp_dashboard_url = string
grafana_istio_mesh_dashboard_url = string
grafana_istio_performance_dashboard_url = string
grafana_istio_service_dashboard_url = string

prometheus_metrics_endpoint = string
})
| `null` | no | | [java\_config](#input\_java\_config) | Configuration object for Java/JMX monitoring |
object({
enable_alerting_rules = bool
enable_recording_rules = bool
enable_dashboards = bool
scrape_sample_limit = number


flux_gitrepository_name = string
flux_gitrepository_url = string
flux_gitrepository_branch = string
flux_kustomization_name = string
flux_kustomization_path = string

grafana_dashboard_url = string

prometheus_metrics_endpoint = string
})
| `null` | no | | [ksm\_config](#input\_ksm\_config) | Kube State metrics configuration |
object({
create_namespace = bool
k8s_namespace = string
helm_chart_name = string
helm_chart_version = string
helm_release_name = string
helm_repo_url = string
helm_settings = map(string)
helm_values = map(any)

scrape_interval = string
scrape_timeout = string
})
|
{
"create_namespace": true,
"helm_chart_name": "kube-state-metrics",
"helm_chart_version": "4.24.0",
"helm_release_name": "kube-state-metrics",
"helm_repo_url": "https://prometheus-community.github.io/helm-charts",
"helm_settings": {},
"helm_values": {},
"k8s_namespace": "kube-system",
"scrape_interval": "60s",
"scrape_timeout": "15s"
}
| no | | [logs\_config](#input\_logs\_config) | Configuration object for logs collection |
object({
cw_log_retention_days = number
})
|
{
"cw_log_retention_days": 90
}
| no | diff --git a/modules/eks-monitoring/locals.tf b/modules/eks-monitoring/locals.tf index 8125703..59183d3 100644 --- a/modules/eks-monitoring/locals.tf +++ b/modules/eks-monitoring/locals.tf @@ -77,4 +77,32 @@ locals { grafana_url = var.grafana_url grafana_dashboard_url = "https://raw.githubusercontent.com/aws-observability/aws-observability-accelerator/main/artifacts/grafana-dashboards/eks/nginx/nginx.json" } + + istio_pattern_config = { + # disabled if options from module are disabled, by default + # can be overriden by providing a config + enable_alerting_rules = var.enable_alerting_rules + enable_recording_rules = var.enable_recording_rules + enable_dashboards = var.enable_dashboards + + scrape_sample_limit = 1000 + + flux_gitrepository_name = "aws-observability-accelerator" + flux_gitrepository_url = "https://github.com/aws-observability/aws-observability-accelerator" + flux_gitrepository_branch = "main" + flux_kustomization_name = "grafana-dashboards-istio" + flux_kustomization_path = "./artifacts/grafana-operator-manifests/eks/istio" + + managed_prometheus_workspace_id = var.managed_prometheus_workspace_id + managed_prometheus_workspace_region = var.managed_prometheus_workspace_region + managed_prometheus_workspace_endpoint = var.managed_prometheus_workspace_endpoint + prometheus_metrics_endpoint = "/metrics" + + grafana_url = var.grafana_url + grafana_istio_cp_dashboard_url = "https://raw.githubusercontent.com/aws-observability/aws-observability-accelerator/main/artifacts/grafana-dashboards/eks/istio/istio-control-plane-dashboard.json" + grafana_istio_mesh_dashboard_url = "https://raw.githubusercontent.com/aws-observability/aws-observability-accelerator/main/artifacts/grafana-dashboards/eks/istio/istio-mesh-dashboard.json" + grafana_istio_performance_dashboard_url = "https://raw.githubusercontent.com/aws-observability/aws-observability-accelerator/main/artifacts/grafana-dashboards/eks/istio/istio-performance-dashboard.json" + grafana_istio_service_dashboard_url = "https://raw.githubusercontent.com/aws-observability/aws-observability-accelerator/main/artifacts/grafana-dashboards/eks/istio/istio-service-dashboard.json" + } + } diff --git a/modules/eks-monitoring/main.tf b/modules/eks-monitoring/main.tf index 3dcfc53..308a04e 100644 --- a/modules/eks-monitoring/main.tf +++ b/modules/eks-monitoring/main.tf @@ -169,6 +169,18 @@ module "helm_addon" { name = "nginxPrometheusMetricsEndpoint" value = try(var.nginx_config.prometheus_metrics_endpoint, local.nginx_pattern_config.prometheus_metrics_endpoint) }, + { + name = "enableIstio" + value = var.enable_istio + }, + { + name = "istioScrapeSampleLimit" + value = try(var.istio_config.scrape_sample_limit, local.istio_pattern_config.scrape_sample_limit) + }, + { + name = "istioPrometheusMetricsEndpoint" + value = try(var.istio_config.prometheus_metrics_endpoint, local.istio_pattern_config.prometheus_metrics_endpoint) + } ] irsa_config = { @@ -202,6 +214,13 @@ module "nginx_monitoring" { pattern_config = coalesce(var.nginx_config, local.nginx_pattern_config) } +module "istio_monitoring" { + source = "./patterns/istio" + count = var.enable_istio ? 1 : 0 + + pattern_config = coalesce(var.istio_config, local.istio_pattern_config) +} + module "fluentbit_logs" { source = "./add-ons/aws-for-fluentbit" count = var.enable_logs ? 1 : 0 diff --git a/modules/eks-monitoring/otel-config/templates/opentelemetrycollector.yaml b/modules/eks-monitoring/otel-config/templates/opentelemetrycollector.yaml index 044c346..850bbca 100644 --- a/modules/eks-monitoring/otel-config/templates/opentelemetrycollector.yaml +++ b/modules/eks-monitoring/otel-config/templates/opentelemetrycollector.yaml @@ -1562,6 +1562,67 @@ spec: action: labeldrop {{ end }} + {{ if .Values.enableIstio }} + - honor_labels: true + job_name: kubernetes-istio + kubernetes_sd_configs: + - role: pod + relabel_configs: + - action: keep + regex: true + source_labels: + - __meta_kubernetes_pod_annotation_prometheus_io_scrape + - action: drop + regex: true + source_labels: + - __meta_kubernetes_pod_annotation_prometheus_io_scrape_slow + - action: replace + regex: (https?) + source_labels: + - __meta_kubernetes_pod_annotation_prometheus_io_scheme + target_label: __scheme__ + - action: replace + regex: (.+) + source_labels: + - __meta_kubernetes_pod_annotation_prometheus_io_path + target_label: __metrics_path__ + - action: replace + regex: (\d+);(([A-Fa-f0-9]{1,4}::?){1,7}[A-Fa-f0-9]{1,4}) + replacement: '[$$2]:$$1' + source_labels: + - __meta_kubernetes_pod_annotation_prometheus_io_port + - __meta_kubernetes_pod_ip + target_label: __address__ + - action: replace + regex: (\d+);((([0-9]+?)(\.|$)){4}) + replacement: $$2:$$1 + source_labels: + - __meta_kubernetes_pod_annotation_prometheus_io_port + - __meta_kubernetes_pod_ip + target_label: __address__ + - action: labelmap + regex: __meta_kubernetes_pod_annotation_prometheus_io_param_(.+) + replacement: __param_$1 + - action: labelmap + regex: __meta_kubernetes_pod_label_(.+) + - action: replace + source_labels: + - __meta_kubernetes_namespace + target_label: namespace + - action: replace + source_labels: + - __meta_kubernetes_pod_name + target_label: pod + - action: keep + source_labels: [ __address__ ] + regex: '.*:15020$$' + - action: drop + regex: Pending|Succeeded|Failed|Completed + source_labels: + - __meta_kubernetes_pod_phase + {{ end }} + + exporters: {{ if .Values.enableTracing }} awsxray: diff --git a/modules/eks-monitoring/otel-config/values.yaml b/modules/eks-monitoring/otel-config/values.yaml index 8d77bc8..64e6998 100644 --- a/modules/eks-monitoring/otel-config/values.yaml +++ b/modules/eks-monitoring/otel-config/values.yaml @@ -24,4 +24,8 @@ enableNginx: ${enable_nginx} nginxScrapeSampleLimit: ${nginx_scrape_sample_limit} nginxPrometheusMetricsEndpoint: ${nginx_prometheus_metrics_endpoint} +enableIstio: ${enable_istio} +istioScrapeSampleLimit: ${istio_scrape_sample_limit} +istioPrometheusMetricsEndpoint: ${istio_prometheus_metrics_endpoint} + adotLoglevel: ${adot_loglevel} diff --git a/modules/eks-monitoring/patterns/istio/README.md b/modules/eks-monitoring/patterns/istio/README.md new file mode 100644 index 0000000..fc9feda --- /dev/null +++ b/modules/eks-monitoring/patterns/istio/README.md @@ -0,0 +1,47 @@ +# Istio patterns module + +Provides monitoring for Istio based workloads with the following resources: + +- AWS Managed Grafana Dashboard and data source +- Alerts and recording rules with AWS Managed Service for Prometheus + + +## Requirements + +| Name | Version | +|------|---------| +| [terraform](#requirement\_terraform) | >= 1.1.0 | +| [aws](#requirement\_aws) | >= 4.0.0 | +| [helm](#requirement\_helm) | >= 2.4.1 | +| [kubectl](#requirement\_kubectl) | >= 1.14 | +| [kubernetes](#requirement\_kubernetes) | >= 2.10 | + +## Providers + +| Name | Version | +|------|---------| +| [aws](#provider\_aws) | >= 4.0.0 | +| [kubectl](#provider\_kubectl) | >= 1.14 | + +## Modules + +No modules. + +## Resources + +| Name | Type | +|------|------| +| [aws_prometheus_rule_group_namespace.alerting_rules](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/resources/prometheus_rule_group_namespace) | resource | +| [aws_prometheus_rule_group_namespace.recording_rules](https://registry.terraform.io/providers/hashicorp/aws/latest/docs/resources/prometheus_rule_group_namespace) | resource | +| [kubectl_manifest.flux_kustomization](https://registry.terraform.io/providers/gavinbunney/kubectl/latest/docs/resources/manifest) | resource | + +## Inputs + +| Name | Description | Type | Default | Required | +|------|-------------|------|---------|:--------:| +| [pattern\_config](#input\_pattern\_config) | Configuration object for ISTIO monitoring |
object({
enable_alerting_rules = bool
enable_recording_rules = bool
scrape_sample_limit = number

enable_recording_rules = bool

enable_dashboards = bool

flux_gitrepository_name = string
flux_gitrepository_url = string
flux_gitrepository_branch = string
flux_kustomization_name = string
flux_kustomization_path = string

managed_prometheus_workspace_id = string
managed_prometheus_workspace_region = string
managed_prometheus_workspace_endpoint = string

grafana_url = string
grafana_istio_cp_dashboard_url = string
grafana_istio_mesh_dashboard_url = string
grafana_istio_performance_dashboard_url = string
grafana_istio_service_dashboard_url = string
})
| n/a | yes | + +## Outputs + +No outputs. + diff --git a/modules/eks-monitoring/patterns/istio/main.tf b/modules/eks-monitoring/patterns/istio/main.tf new file mode 100644 index 0000000..8e0f48e --- /dev/null +++ b/modules/eks-monitoring/patterns/istio/main.tf @@ -0,0 +1,217 @@ +resource "aws_prometheus_rule_group_namespace" "recording_rules" { + count = var.pattern_config.enable_recording_rules ? 1 : 0 + + name = "accelerator-istio-rules" + workspace_id = var.pattern_config.managed_prometheus_workspace_id + data = < + absent(istio_requests_total{destination_service_namespace=~"service-graph.*",reporter="source",source_workload="istio-ingressgateway"})==1 + for: 5m + - alert: IstioMetricsMissing + annotations: + summary: 'Istio Metrics missing' + description: '[Critical]: Check prometheus deployment or whether the prometheus filters are applied correctly' + expr: > + absent(istio_request_total)==1 or absent(istio_request_duration_milliseconds_bucket)==1 + for: 5m + - name: "istio.workload.alerting-rules" + rules: + - alert: HTTP5xxRateHigh + annotations: + summary: '5xx rate too high' + description: 'The HTTP 5xx errors rate higher than 0.05 in 5 mins' + expr: > + sum(irate(istio_requests_total{reporter="destination", response_code=~"5.*"}[5m])) / sum(irate(istio_requests_total{reporter="destination"}[5m])) > 0.05 + for: 5m + - alert: WorkloadLatencyP99High + expr: histogram_quantile(0.99, sum(irate(istio_request_duration_milliseconds_bucket{source_workload=~"svc.*"}[5m])) by (source_workload,namespace, le)) > 160 + for: 10m + annotations: + description: 'The workload request latency P99 > 160ms ' + message: "Request duration has slowed down for workload: {{`{{$labels.source_workload}}`}} in namespace: {{`{{$labels.namespace}}`}}. Response duration is {{`{{$value}}`}} milliseconds" + - alert: IngressLatencyP99High + expr: histogram_quantile(0.99, sum(irate(istio_request_duration_milliseconds_bucket{source_workload=~"istio.*"}[5m])) by (source_workload,namespace, le)) > 250 + for: 10m + annotations: + description: 'The ingress latency P99 > 250ms ' + message: "Request duration has slowed down for ingress: {{`{{$labels.source_workload}}`}} in namespace: {{`{{$labels.namespace}}`}}. Response duration is {{`{{$value}}`}} milliseconds" + - name: "istio.infra.alerting-rules" + rules: + - alert: ProxyContainerCPUUsageHigh + expr: (sum(rate(container_cpu_usage_seconds_total{namespace!="kube-system", container=~"istio-proxy", namespace!=""}[5m])) BY (namespace, pod, container) * 100) > 80 + for: 5m + annotations: + summary: "Proxy Container CPU usage (namespace {{ $labels.namespace }}) (pod {{ $labels.pod }}) (container {{ $labels.container }}) VALUE = {{ $value }}\n" + description: "Proxy Container CPU usage is above 80%" + - alert: ProxyContainerMemoryUsageHigh + expr: (sum(container_memory_working_set_bytes{namespace!="kube-system", container=~"istio-proxy", namespace!=""}) BY (container, pod, namespace) / (sum(container_spec_memory_limit_bytes{namespace!="kube-system", container!="POD"}) BY (container, pod, namespace) > 0)* 100) > 80 + for: 5m + annotations: + summary: "Proxy Container Memory usage (namespace {{ $labels.namespace }}) (pod {{ $labels.pod }}) (container {{ $labels.container }}) VALUE = {{ $value }}\n" + description: "Proxy Container Memory usage is above 80%" + - alert: IngressMemoryUsageIncreaseRateHigh + expr: avg(deriv(container_memory_working_set_bytes{container=~"istio-proxy",namespace="istio-system"}[60m])) > 200 + for: 180m + annotations: + summary: "Ingress proxy Memory change rate, VALUE = {{ $value }}\n" + description: "Ingress proxy Memory Usage increases more than 200 Bytes/sec" + - alert: IstiodContainerCPUUsageHigh + expr: (sum(rate(container_cpu_usage_seconds_total{namespace="istio-system", container="discovery"}[5m])) BY (pod) * 100) > 80 + for: 5m + annotations: + summary: "Istiod Container CPU usage (namespace {{ $labels.namespace }}) (pod {{ $labels.pod }}) (container {{ $labels.container }}) VALUE = {{ $value }}\n" + description: "Isitod Container CPU usage is above 80%" + - alert: IstiodMemoryUsageHigh + expr: (sum(container_memory_working_set_bytes{namespace="istio-system", container="discovery"}) BY (pod) / (sum(container_spec_memory_limit_bytes{namespace="istio-system", container="discovery"}) BY (pod) > 0)* 100) > 80 + for: 5m + annotations: + summary: "Istiod Container Memory usage (namespace {{ $labels.namespace }}) (pod {{ $labels.pod }}) (container {{ $labels.container }}) VALUE = {{ $value }}\n" + description: "Istiod Container Memory usage is above 80%" + - alert: IstiodMemoryUsageIncreaseRateHigh + expr: sum(deriv(container_memory_working_set_bytes{namespace="istio-system",pod=~"istiod-.*"}[60m])) > 1000 + for: 300m + annotations: + summary: "Istiod Container Memory usage increase rate high, VALUE = {{ $value }}\n" + description: "Istiod Container Memory usage increases more than 1k Bytes/sec" + - name: "istio.controlplane.alerting-rules" + rules: + - alert: IstiodxdsPushErrorsHigh + annotations: + summary: 'istiod push errors is too high' + description: 'istiod push error rate is higher than 0.05' + expr: > + sum(irate(pilot_xds_push_errors{app="istiod"}[5m])) / sum(irate(pilot_xds_pushes{app="istiod"}[5m])) > 0.05 + for: 5m + - alert: IstiodxdsRejectHigh + annotations: + summary: 'istiod rejects rate is too high' + description: 'istiod rejects rate is higher than 0.05' + expr: > + sum(irate(pilot_total_xds_rejects{app="istiod"}[5m])) / sum(irate(pilot_xds_pushes{app="istiod"}[5m])) > 0.05 + for: 5m + - alert: IstiodContainerNotReady + annotations: + summary: 'istiod container not ready' + description: 'container: discovery not running' + expr: > + kube_pod_container_status_running{namespace="istio-system", container="discovery", component=""} == 0 + for: 5m + - alert: IstiodUnavailableReplica + annotations: + summary: 'Istiod unavailable pod' + description: 'Istiod unavailable replica > 0' + expr: > + kube_deployment_status_replicas_unavailable{deployment="istiod", component=""} > 0 + for: 5m + - alert: Ingress200RateLow + annotations: + summary: 'ingress gateway 200 rate drops' + description: 'The expected rate is 100 per ns, the limit is set based on 15ns' + expr: > + sum(rate(istio_requests_total{reporter="source", source_workload="istio-ingressgateway",response_code="200",destination_service_namespace=~"service-graph.*"}[5m])) < 1490 + for: 30m + EOF +} + +resource "kubectl_manifest" "flux_kustomization" { + count = var.pattern_config.enable_dashboards ? 1 : 0 + + yaml_body = <