* changes to just OTEL kube-proxy scrape job
* agree with koffir@ to implement kube-proxy by default
* re-added KubeProxyDown Prometheus alert rule
* added kube-proxy URL related stuff
* corrected kube-proxy stuff according to new structure
* Use SEMVER for flux now
* Run pre-commit
* Bump OTEL chart version
* Substitute URL path with SEMVER
* Run pre-commit
* Added kube-proxy config object
* Run pre-commit
* Use kube-proxy config object for dashboards
---------
Co-authored-by: Jens-Uwe Walther <waltju@amazon.com>
* adot-container-insight-tf-code by Rajat Omar
* Documentation and naming convention added
* PR review fixes
* PR CI pipeline fixes
* pr ci run fixes
* added the variables mentioned in ci build
* ci variable fixes
* changed the module name
---------
Co-authored-by: Omar <merajat@3c0630162a5a.ant.amazon.com>
* added workqueue related and apiserver_storage_db_total_size_in_bytes (available since K8s/EKS v1.26+) metrics into kube-admin scrape job
* chnanged OTEL scrape config and recording rules file to make apiserver Grafana dashboards working
* added APISERVER Grafana dashboards into variables, Flux kustomization
* added original kube-prom-stack kube-apiserver scrape config into OTEL
* Revert "added original kube-prom-stack kube-apiserver scrape config into OTEL"
because this scrape config does not work :-(
This reverts commit 715db895c656af09430e3d6ada2087bd1828413f.
* added API server troubleshoting dashboard
* removed empty line for clarity
* make API serve rmonitoring default to true but can be disabled as well
* Update eks-apiserver.md
* updated eks-monitoring/README.md according to pre-commit
---------
Co-authored-by: Jens-Uwe Walther <waltju@amazon.com>
* Typo
* Remove Grafana provider
* Temp: move dashbaords to gitOps
* Move external labels to resource attributes
* Avoid DDoS with using 0.0.0.0
* Pre-commit
* Transition in two steps
Will need to remove provider in a separate version to provide a transition path as removing this will break terraform and leave orphans in the state
* Move patterns' dashboards creation to gitOps
Standardize config objects for patterns as well
* Pre-commit
* Create AMP dashboard from external source with Grafana provider
* Fix deprecated option
* Fix Flux requirements
* Run pre-commit
* Update example with operator
* Cleanup examples
* Update multicluster example
* Update multicluster example
* Drop dead variable
* Update docs
* Change GitOps branch name
* Update docs
* Replacing Secrets Manager to SSM to store Grafana API Key (#178)
* Fixing SSM
* Fixing SSM
* Replacing Secrets Manager with SSM
* Replacing Secrets Manager with SSM
* Update architecture diagram
* Update architecture diagram
* Update README.md
* Update index.md
* Fixing Grafana Operator Version
* Fix multicluster example
* Update docs
---------
Co-authored-by: Ela AWS <51791117+elamaran11@users.noreply.github.com>
Co-authored-by: Elamaran Shanmugam <elamaran.shan@gmail.com>
* Fixing breaking changes for path and dash urls
* Fixing breaking changes for path and dash urls
* Fixing breaking changes for path and dash urls
* Fixing breaking changes for path and dash urls
* Grafana With GitOps Feature
* Grafana With GitOps Feature
* Grafana With GitOps Feature
* Fix setup logs retention policy (#169)
* Fixing GitOps Repo
* Commenting out the NodeExp Dash
* Commenting out the NodeExp Dash
* Adding all Grafana Dashboards
* Adding all Grafana Dashboards
* Fixing Grafana Operator Version and cleaning full boards
* Fixing Grafana Operator Version and cleaning full boards
* Fixing Grafana Operator Version and cleaning full boards
* Fixing Grafana Operator Version and cleaning full boards
* Fixing Grafana Operator Version and cleaning full boards and PR Issues
* Fixing Grafana Operator Version and cleaning full boards and PR Issues
* Fixing Grafana Operator Version and cleaning full boards and PR Issues
* Fixing Grafana Operator Version and cleaning full boards and PR Issues
* Fixing Grafana Operator Version and cleaning full boards and PR Issues
* Fixing Grafana Operator Version and cleaning full boards and PR Issues
---------
Co-authored-by: Rodrigue Koffi <bonclay7@users.noreply.github.com>
* Made the cluster variable visible in dashboards
* Made the repo reusable with multiple EKS clusters
* Relabled the input variable from create_grafana_data_source to create_prometheus_data_source
* Import and customize fluenbit add-on
* Enable fluent bit logs
* Bump helm addon version
* Dropping account id as it seems to create scraping errors
* Create separate log groups per namespace
* Apply pre-commit
* Remove conflicting global label
* Add config object for logs
* Enable logs in examples
* Add logs docs
* Fix broken link
* Add screenshots
* Update docs
* Typos
* Validation added to cluster_name variable
* Update ref to eks-cluster-with-vpc on v4.13.1 (#121)
* Added cluster_name validation with message (#118)
* Error message typo fixed to follow the standards.
* Adding requested new line at the end
* Remove duplicate exposed ports in ADOT 0.61+
* Sync ADOT collector to ADOT addon version
* Bump helm version
* Provide configuration options for tracing
* Update default helm chart version
Update default helm chart version for node-exporter and kube-state-metrics
* update README
* revert node-exporter chart version
* Init tracing support
* Update xray exporter
* Remove duplicate metrics only on daemon set
* Add support for custom metrics collection
* Remove experimental feature
* Rename config items
* Update requirements for tracing support