Move all dashboards to GitOps (#175)

* Typo

* Remove Grafana provider

* Temp: move dashbaords to gitOps

* Move external labels to resource attributes

* Avoid DDoS with using 0.0.0.0

* Pre-commit

* Transition in two steps

Will need to remove provider in a separate version to provide a transition path as removing this will break terraform and leave orphans in the state

* Move patterns' dashboards creation to gitOps

Standardize config objects for patterns as well

* Pre-commit

* Create AMP dashboard from external source with Grafana provider

* Fix deprecated option

* Fix Flux requirements

* Run pre-commit

* Update example with operator

* Cleanup examples

* Update multicluster example

* Update multicluster example

* Drop dead variable

* Update docs

* Change GitOps branch name

* Update docs

* Replacing Secrets Manager to SSM to store Grafana API Key (#178)

* Fixing SSM

* Fixing SSM

* Replacing Secrets Manager with SSM

* Replacing Secrets Manager with SSM

* Update architecture diagram

* Update architecture diagram

* Update README.md

* Update index.md

* Fixing Grafana Operator Version

* Fix multicluster example

* Update docs

---------

Co-authored-by: Ela AWS <51791117+elamaran11@users.noreply.github.com>
Co-authored-by: Elamaran Shanmugam <elamaran.shan@gmail.com>
This commit is contained in:
Rodrigue Koffi
2023-06-12 18:00:42 +02:00
committed by GitHub
parent c5e4c0c718
commit fa38a90efc
46 changed files with 363 additions and 4462 deletions
+13 -11
View File
@@ -39,22 +39,24 @@ The grafana-operator is a Kubernetes operator built to help you manage your Graf
GitOps is a way of managing application and infrastructure deployment so that the whole system is described declaratively in a Git repository. It is an operational model that offers you the ability to manage the state of multiple Kubernetes clusters leveraging the best practices of version control, immutable artifacts, and automation. Flux is a declarative, GitOps-based continuous delivery tool that can be integrated into any CI/CD pipeline. It gives users the flexibility of choosing their Git provider (GitHub, GitLab, BitBucket). Now, with grafana-operator supporting the management of external Grafana instances such as Amazon Managed Grafana, operations personas can use GitOps mechanisms using CNCF projects such as Flux to create and manage the lifecycle of resources in Amazon Managed Grafana.
We have setup a [GitRepository](https://fluxcd.io/flux/components/source/gitrepositories/) and [Kustomization](https://fluxcd.io/flux/components/kustomize/kustomization/) using flux to sync our GitHub Repository to add Grafana Datasources, folder and Dashboards to Amazon Managed Grafana using Grafana Operator. GitRepository defines a Source to produce an Artifact for a Git repository revision. Kustomization defines a pipeline for fetching, decrypting, building, validating and applying Kustomize overlays or plain Kubernetes manifests. we are also using [Flux Post build variable substitution](https://fluxcd.io/flux/components/kustomize/kustomization/#post-build-variable-substitution) to dynamically render variables such as AMG_AWS_REGION, AMP_ENDPOINT_URL, AMG_ENDPOINT_URL,GRAFANA_NODEEXP_DASH_URL on the YAML manifests during deployment time to avoid hardcoding on the YAML manifests stored in Git repo.
We have setup a [GitRepository](https://fluxcd.io/flux/components/source/gitrepositories/) and [Kustomization](https://fluxcd.io/flux/components/kustomize/kustomization/) using Flux to sync our GitHub Repository to add Grafana Datasources, folder and Dashboards to Amazon Managed Grafana using Grafana Operator. GitRepository defines a Source to produce an Artifact for a Git repository revision. Kustomization defines a pipeline for fetching, decrypting, building, validating and applying Kustomize overlays or plain Kubernetes manifests. we are also using [Flux Post build variable substitution](https://fluxcd.io/flux/components/kustomize/kustomization/#post-build-variable-substitution) to dynamically render variables such as AMG_AWS_REGION, AMP_ENDPOINT_URL, AMG_ENDPOINT_URL,GRAFANA_NODEEXP_DASH_URL on the YAML manifests during deployment time to avoid hardcoding on the YAML manifests stored in Git repo.
We have placed our declarative code snippet to create an Amazon Managed Service For Promethes datasource and Grafana Dashboard in Amazon Managed Grafana in our [AWS Observabiity Accelerator GitHub Repository](https://github.com/aws-observability/aws-observability-accelerator/tree/main/artifacts/grafana-operator-manifests). We have setup a GitRepository to point to the AWS Observabiity Accelerator GitHub Repository and `Kustomization` for flux to sync Git Repository with artifacts in `./artifacts/grafana-operator-manifests` path in the AWS Observabiity Accelerator GitHub Repository. You can use this extension of our solution to point your own Kubernetes manifests to create Grafana Datasources and personified Grafana Dashboards of your choice using GitOps with Grafana Operator and Flux in Kubernetes native way with altering and redeploying this solution for changes to Grafana resources.
We have placed our declarative code snippet to create an Amazon Managed Service For Promethes datasource and Grafana Dashboard in Amazon Managed Grafana in our [AWS Observabiity Accelerator GitHub Repository](https://github.com/aws-observability/aws-observability-accelerator). We have setup a GitRepository to point to the AWS Observabiity Accelerator GitHub Repository and `Kustomization` for flux to sync Git Repository with artifacts in `./artifacts/grafana-operator-manifests/*` path in the AWS Observabiity Accelerator GitHub Repository. You can use this extension of our solution to point your own Kubernetes manifests to create Grafana Datasources and personified Grafana Dashboards of your choice using GitOps with Grafana Operator and Flux in Kubernetes native way with altering and redeploying this solution for changes to Grafana resources.
## v2.x changes
## Release notes
v2.x [releases](https://github.com/aws-observability/terraform-aws-observability-accelerator/releases) introduce
couple of breaking changes compared to previous versions:
We encourage you to use our [release versions](https://github.com/aws-observability/terraform-aws-observability-accelerator/releases)
as much as possible to avoid breaking changes when deploying Terraform modules. You can
read also our change log on the releases page. Here's an example of using a fixed version:
```hcl
module "eks_monitoring" {
source = "github.com/aws-observability/terraform-aws-observability-accelerator//modules/managed-prometheus-monitoring?ref=v2.5.0"
}
```
- `modules/workloads/infra` module moves to `modules/eks-monitoring`
- EKS configuration options moves from the base module to the `eks-monitoring` module
- EKS workload modules **java,nginx** merge into `eks-monitoring` as configuration options (patterns),
see [examples](https://github.com/aws-observability/terraform-aws-observability-accelerator/tree/main/examples)
- Examples have been updated to reflect these changes
## Base module
@@ -138,4 +140,4 @@ classDiagram
If you are new to AWS Observability services, or want to dive deeper into them,
check our [One Observability Workshop](https://catalog.workshops.aws/observability/)
for a hands-on experience in a self-paced environement or at an AWS venue.
for a hands-on experience in a self-paced environment or at an AWS venue.
+38 -22
View File
@@ -111,26 +111,41 @@ terraform apply
## Visualization
#### 1. Prometheus data source on Grafana
Make sure to open the link in the output. After a successful deployment, this will open
the Prometheus data source configuration on Grafana.
Click `Save & test` and you should see a notification confirming that the Amazon Managed Service for Prometheus workspace is ready to be used on Grafana.
```bash
terraform output grafana_prometheus_datasource_test
```
#### 2. Grafana dashboards
Go to the Dashboards panel of your Grafana workspace. You should see a list of dashboards under the `Observability Accelerator Dashboards`
#### 1. Grafana dashboards
Login to your Grafana workspace and navigate to the Dashboards panel. You should see a list of dashboards under the `Observability Accelerator Dashboards`
<img width="1540" alt="image" src="https://user-images.githubusercontent.com/10175027/190000716-29e16698-7c90-49d6-8c37-79ca1790e2cc.png">
Open a specific dashboard and you should be able to view its visualization
<img width="2056" alt="cluster headlines" src="https://user-images.githubusercontent.com/10175027/199110753-9bc7a9b7-1b45-4598-89d3-32980154080e.png">
With v2.5 and above, the dashboards are managed with a Grafana Operator running in your cluster.
From the cluster to view all dashboards as Kubernetes objects, run
```console
kubectl get grafanadashboards -A
NAMESPACE NAME AGE
grafana-operator cluster-grafanadashboard 138m
grafana-operator java-grafanadashboard 143m
grafana-operator kubelet-grafanadashboard 13h
grafana-operator namespace-workloads-grafanadashboard 13h
grafana-operator nginx-grafanadashboard 134m
grafana-operator node-exporter-grafanadashboard 13h
grafana-operator nodes-grafanadashboard 13h
grafana-operator workloads-grafanadashboard 13h
```
You can inspect more details per dashboard using this command
```console
kubectl describe grafanadashboards cluster-grafanadashboard -n grafana-operator
```
Grafana Operator and Flux always work together to synchronize your dashboards with Git.
If you delete your dashboards by accident, they will be re-provisioned automatically.
#### 3. Amazon Managed Service for Prometheus rules and alerts
Open the Amazon Managed Service for Prometheus console and view the details of your workspace. Under the `Rules management` tab, you should find new rules deployed.
@@ -216,21 +231,22 @@ export GO_AMG_API_KEY=$(aws grafana create-workspace-api-key \
--output text)
```
- Next, lets grab the Grafana API key secret name from AWS Secrets Manager. The keyname should start with `terraform-..`
```bash
aws secretsmanager list-secrets
```
- Finally, update the Grafana API key secret in AWS Secrets Manager using the above new Grafana API key:
```bash
aws secretsmanager update-secret \
--secret-id <Your Secret Name> \
--secret-string "{\"GF_SECURITY_ADMIN_APIKEY\": \"${GO_AMG_API_KEY}\"}" \
aws aws ssm put-parameter \
--name "/terraform-accelerator/grafana-api-key" \
--type "SecureString" \
--value "{\"GF_SECURITY_ADMIN_APIKEY\": \"${GO_AMG_API_KEY}\"}" \
--region <Your AWS Region>
```
- If the issue persists, you can force the synchronization by deleting the `externalsecret` Kubernetes object.
```bash
kubectl delete externalsecret/external-secrets-sm -n grafana-operator
```
### 2. Upgrade from 2.1.0 or earlier
When you upgrade the eks-monitoring module from v2.1.0 or earlier, the following error may occur.
+1 -1
View File
@@ -32,7 +32,7 @@ Make sure to refresh your temporary Grafana API key
```bash
export TF_VAR_managed_grafana_workspace_id=g-xxx
export TF_VAR_grafana_api_key=`aws grafana create-workspace-api-key --key-name "observability-accelerator-$(date +%s)" --key-role ADMIN --seconds-to-live 1200 --workspace-id $TF_VAR_managed_grafana_workspace_id --query key --output text`
export TF_VAR_grafana_api_key=`aws grafana create-workspace-api-key --key-name "observability-accelerator-$(date +%s)" --key-role ADMIN --seconds-to-live 7200 --workspace-id $TF_VAR_managed_grafana_workspace_id --query key --output text`
```
## Deploy
+4 -4
View File
@@ -11,7 +11,7 @@ Using the example [eks-cluster-with-vpc](https://aws-observability.github.io/ter
1. `eks-cluster-1`
2. `eks-cluster-2`
#### 2. Amazon Managed Serivce for Prometheus (AMP) workspace
#### 2. Amazon Managed Service for Prometheus (AMP) workspace
We recommend that you create a new AMP workspace. To do that you can run the following command.
@@ -48,7 +48,7 @@ Ensure you have the following necessary IAM permissions
* `grafana.DeleteWorkspaceApiKey`
```sh
export TF_VAR_grafana_api_key=`aws grafana create-workspace-api-key --key-name "observability-accelerator-$(date +%s)" --key-role ADMIN --seconds-to-live 1200 --workspace-id $TF_VAR_managed_grafana_workspace_id --query key --output text`
export TF_VAR_grafana_api_key=`aws grafana create-workspace-api-key --key-name "observability-accelerator-$(date +%s)" --key-role ADMIN --seconds-to-live 7200 --workspace-id $TF_VAR_managed_grafana_workspace_id --query key --output text`
```
## Setup
@@ -70,8 +70,8 @@ Verify by looking at the file `variables.tf` that there are two EKS clusters tar
The difference in deployment between these clusters is that Terraform, when setting up the EKS cluster behind variable `eks_cluster_1_id` for observability, also sets up:
* Dashboard folder and files in `AMG`
* Prometheus and Java, alerting and recording rules in `AMP`
* Dashboard folder and files in Amazon Managed Grafana
* Prometheus and Java, alerting and recording rules in Amazon Managed Service for Prometheus
!!! warning
To override the defaults, create a `terraform.tfvars` and change the default values of the variables.
+2 -2
View File
@@ -14,7 +14,7 @@ your custom applications.
You also can monitor your Amazon Managed Service for Prometheus workspaces ingestion,
costs, active series with [this module](https://aws-observability.github.io/terraform-aws-observability-accelerator/workloads/managed-prometheus/).
<img width="1501" alt="image" src="images/dark-o11y-accelerator-amp-xray.png">
![image](https://github.com/aws-observability/terraform-aws-observability-accelerator/assets/10175027/e83f8709-f754-4192-90f2-e3de96d2e26c)
## Getting started
@@ -26,7 +26,7 @@ traces collection, dashboards and alerts for monitoring:
- Java/JMX workloads (running on Amazon EKS)
- Amazon Managed Service for Prometheus workspaces with Amazon CloudWatch
- [Grafana Operator](https://github.com/grafana-operator/grafana-operator) and [Flux CD](https://fluxcd.io/) to manage Grafana contents (AWS data sources, Grafana Dashboards) with GitOps
- External Secrets Operator to retrieve and Sync the Grafana API keys
- External Secrets Operator to retrieve and sync the Grafana API keys
These modules can be directly configured in your existing Terraform
configurations or ready to be deployed in our packaged