* changes to just OTEL kube-proxy scrape job
* agree with koffir@ to implement kube-proxy by default
* re-added KubeProxyDown Prometheus alert rule
* added kube-proxy URL related stuff
* corrected kube-proxy stuff according to new structure
* Use SEMVER for flux now
* Run pre-commit
* Bump OTEL chart version
* Substitute URL path with SEMVER
* Run pre-commit
* Added kube-proxy config object
* Run pre-commit
* Use kube-proxy config object for dashboards
---------
Co-authored-by: Jens-Uwe Walther <waltju@amazon.com>
* added workqueue related and apiserver_storage_db_total_size_in_bytes (available since K8s/EKS v1.26+) metrics into kube-admin scrape job
* chnanged OTEL scrape config and recording rules file to make apiserver Grafana dashboards working
* added APISERVER Grafana dashboards into variables, Flux kustomization
* added original kube-prom-stack kube-apiserver scrape config into OTEL
* Revert "added original kube-prom-stack kube-apiserver scrape config into OTEL"
because this scrape config does not work :-(
This reverts commit 715db895c656af09430e3d6ada2087bd1828413f.
* added API server troubleshoting dashboard
* removed empty line for clarity
* make API serve rmonitoring default to true but can be disabled as well
* Update eks-apiserver.md
* updated eks-monitoring/README.md according to pre-commit
---------
Co-authored-by: Jens-Uwe Walther <waltju@amazon.com>
* Typo
* Remove Grafana provider
* Temp: move dashbaords to gitOps
* Move external labels to resource attributes
* Avoid DDoS with using 0.0.0.0
* Pre-commit
* Transition in two steps
Will need to remove provider in a separate version to provide a transition path as removing this will break terraform and leave orphans in the state
* Move patterns' dashboards creation to gitOps
Standardize config objects for patterns as well
* Pre-commit
* Create AMP dashboard from external source with Grafana provider
* Fix deprecated option
* Fix Flux requirements
* Run pre-commit
* Update example with operator
* Cleanup examples
* Update multicluster example
* Update multicluster example
* Drop dead variable
* Update docs
* Change GitOps branch name
* Update docs
* Replacing Secrets Manager to SSM to store Grafana API Key (#178)
* Fixing SSM
* Fixing SSM
* Replacing Secrets Manager with SSM
* Replacing Secrets Manager with SSM
* Update architecture diagram
* Update architecture diagram
* Update README.md
* Update index.md
* Fixing Grafana Operator Version
* Fix multicluster example
* Update docs
---------
Co-authored-by: Ela AWS <51791117+elamaran11@users.noreply.github.com>
Co-authored-by: Elamaran Shanmugam <elamaran.shan@gmail.com>
* Fixing breaking changes for path and dash urls
* Fixing breaking changes for path and dash urls
* Fixing breaking changes for path and dash urls
* Fixing breaking changes for path and dash urls
* Grafana With GitOps Feature
* Grafana With GitOps Feature
* Grafana With GitOps Feature
* Fix setup logs retention policy (#169)
* Fixing GitOps Repo
* Commenting out the NodeExp Dash
* Commenting out the NodeExp Dash
* Adding all Grafana Dashboards
* Adding all Grafana Dashboards
* Fixing Grafana Operator Version and cleaning full boards
* Fixing Grafana Operator Version and cleaning full boards
* Fixing Grafana Operator Version and cleaning full boards
* Fixing Grafana Operator Version and cleaning full boards
* Fixing Grafana Operator Version and cleaning full boards and PR Issues
* Fixing Grafana Operator Version and cleaning full boards and PR Issues
* Fixing Grafana Operator Version and cleaning full boards and PR Issues
* Fixing Grafana Operator Version and cleaning full boards and PR Issues
* Fixing Grafana Operator Version and cleaning full boards and PR Issues
* Fixing Grafana Operator Version and cleaning full boards and PR Issues
---------
Co-authored-by: Rodrigue Koffi <bonclay7@users.noreply.github.com>
* Made the cluster variable visible in dashboards
* Made the repo reusable with multiple EKS clusters
* Relabled the input variable from create_grafana_data_source to create_prometheus_data_source
* Import and customize fluenbit add-on
* Enable fluent bit logs
* Bump helm addon version
* Dropping account id as it seems to create scraping errors
* Create separate log groups per namespace
* Apply pre-commit
* Remove conflicting global label
* Add config object for logs
* Enable logs in examples
* Add logs docs
* Fix broken link
* Add screenshots
* Update docs
* Typos