Production GKE Troubleshooting Toolkit
Overview
When supporting production-grade Google Kubernetes Engine (GKE) clusters across multiple time zones, incident triage consistency becomes a major challenge. Engineers often faced fragmented approaches to debugging, with varying levels of visibility into logs, resource metrics, and node health. To solve this, I led the design and development of a custom troubleshooting toolkit, built using Python and Bash, to accelerate root cause identification and ensure standardized diagnostics across teams.
🎯 Objectives
Build reusable tools to automate first-level diagnostics for GKE clusters.
Reduce the time and complexity involved in parsing logs and checking node status.
Enable global support teams to follow a consistent triage workflow, regardless of timezone or experience level.
Improve documentation of incident handling and RCA data collection.
🔧 What I Built
Log Parsing Script (Python):
Parses container and system logs from
kubectl logs,journalctl, and GCP operations logging APIs.Supports keyword-based error filtering, regex highlighting, and time-range scoping.
Node Health Checker (Bash):
Automates collection of
kubectl describe node, disk I/O stats, kubelet status, and pod distribution checks.Validates readiness and taints across nodes and flags underutilized or failing nodes.
Resource Usage Analyzer (Python/Bash Hybrid):
Gathers CPU/memory requests, limits, and actual usage across all namespaces.
Identifies overprovisioned pods and resource bottlenecks using data from
kubectl topand GCP metrics.
Unified Diagnostic Report Generator:
Outputs results from all tools into a single JSON or HTML report.
Can be shared across Slack/Email during P1/P2 incident calls.
📈 Impact
✅ Reduced average incident triage time by ~40%, especially for after-hours/on-call escalations.
✅ Enabled faster onboarding of new support engineers by providing structured debug steps and tool outputs.
✅ Helped surface recurring issues, feeding directly into SRE-led problem management and platform improvements.
💡 Lessons Learned
Automation tools for support are not just time-savers — they enable knowledge transfer and reduce cognitive load during high-pressure scenarios.
Visibility into “what normal looks like” is critical; baseline outputs were included to help teams compare incident behavior to healthy states.
Collaboration with SREs and Platform Engineers ensured the toolkit fit naturally into existing workflows (Cloud Logging, PagerDuty, etc.).
🚀 Next Steps
Integrate toolkit into a central CLI plugin or internal web UI for broader adoption.
Add support for Anthos clusters and hybrid/on-prem GKE monitoring.
Explore integration with Stackdriver APIs and BigQuery logs for deeper analysis.



