* added workqueue related and apiserver_storage_db_total_size_in_bytes (available since K8s/EKS v1.26+) metrics into kube-admin scrape job * chnanged OTEL scrape config and recording rules file to make apiserver Grafana dashboards working * added APISERVER Grafana dashboards into variables, Flux kustomization * added original kube-prom-stack kube-apiserver scrape config into OTEL * Revert "added original kube-prom-stack kube-apiserver scrape config into OTEL" because this scrape config does not work :-( This reverts commit 715db895c656af09430e3d6ada2087bd1828413f. * added API server troubleshoting dashboard * removed empty line for clarity * make API serve rmonitoring default to true but can be disabled as well * Update eks-apiserver.md * updated eks-monitoring/README.md according to pre-commit --------- Co-authored-by: Jens-Uwe Walther <waltju@amazon.com>
1.6 KiB
Monitoring EKS API server
AWS Distro of OpenTelemetry enables EKS API server monitoring by default and provides three Grafana dashboards:
Kube-apiserver (basic)
The basic dashboard shows metrics recommended in EKS Best Practices Guides - Monitor Control Plane Metrics and provides request rate and latency for API server, latency for ETCD server and overall workqueue sercice time and latency. It allows a drill-down per API server.
Kube-apiserver (advanced)
The advanced dashboard is derived from kube-prometheus-stack "Kubernetes / API server" dashboard and provides a detailed metrics drill-down for example per READ and WRITE operations per component (like deployments, configmaps etc.).
Kube-apiserver (troubleshooting)
This dashboards can be used to troubleshoot API server problems like latency, errors etc.
A detailed description for usage and background information regarding the dashboard can be found in AWS Containers blog post Troubleshooting Amazon EKS API servers with Prometheus.