Architect Scalable Kubernetes Platforms with GitOps
A Kubernetes architect skill for platform design, GitOps, service mesh, multi-tenancy, and cost-optimized cluster operations.
Why it matters
Design and implement robust, scalable, and secure Kubernetes platforms leveraging modern GitOps workflows. Optimize for cost, reliability, and developer experience across multi-cloud and on-premises environments.
Outcomes
What it gets done
Design Kubernetes platform architecture and multi-cluster strategies.
Implement GitOps workflows for continuous delivery and progressive rollouts.
Define and enforce security, multi-tenancy, and compliance patterns.
Optimize Kubernetes clusters for performance, cost, and reliability.
Install
Add it to your toolbox
Run in your project directory:
curl -fsSL https://spark.entire.vc/get/ag-kubernetes-architect | bash Overview
Kubernetes Architect
A Kubernetes architect skill for enterprise platform design, GitOps workflows, service mesh, security, multi-tenancy, and cost-optimized cluster operations. Use it for platform-level Kubernetes architecture and GitOps design, not for local dev clusters, single-node setups, or application-code troubleshooting.
What it does
This is a Kubernetes architect skill specializing in cloud-native infrastructure, GitOps workflows, and enterprise container orchestration at scale, covering managed Kubernetes (EKS, AKS, GKE), enterprise platforms (OpenShift, Rancher, Tanzu), and self-managed clusters (kubeadm, kops, kubespray, air-gapped deployments). GitOps coverage spans ArgoCD, Flux v2, Jenkins X, and Tekton, following the CNCF OpenGitOps principles (declarative, versioned and immutable, pulled automatically, continuously reconciled) plus progressive delivery via Argo Rollouts or Flagger for canary and blue/green strategies. Infrastructure-as-code coverage spans Helm 3.x, Kustomize, Jsonnet, cdk8s, and policy-as-code via OPA, Gatekeeper, and Kyverno admission controllers.
When to use - and when NOT to
Use this skill when designing Kubernetes platform architecture or a multi-cluster strategy, implementing GitOps workflows and progressive delivery, planning service mesh/security/multi-tenancy patterns, or improving reliability, cost, or developer experience in Kubernetes. Security coverage spans Pod Security Standards, network policies, runtime security tools (Falco, Sysdig, Aqua), and supply-chain security (SLSA, Sigstore, SBOM generation). Service mesh coverage spans Istio, Linkerd, Cilium (eBPF-based), Consul Connect, and the Gateway API. Don't use it if you only need a local dev cluster or single-node setup, are troubleshooting application code without platform changes, or aren't using Kubernetes at all. Its safety constraint is explicit: avoid production changes without approvals and rollback plans, and test policy changes and admission controls in staging first.
Inputs and outputs
Input is workload requirements, compliance needs, and scale targets; output follows a nine-step response approach: assess workload requirements, design the Kubernetes architecture for scale, implement GitOps workflows with proper repository structure, configure security policies (Pod Security Standards, network policies), set up an observability stack (Prometheus, Grafana, OpenTelemetry), plan autoscaling and resource management, consider multi-tenancy and namespace isolation, optimize cost (KubeCost, OpenCost, right-sizing, spot instances), and document the platform with operational procedures. Disaster-recovery output covers Velero backups, multi-region active-active/active-passive deployment, and RTO/RPO planning.
Who it's for
Platform engineers and architects designing or scaling enterprise Kubernetes platforms across any major provider or on-premises, who want proven GitOps, security, multi-tenancy, and cost-optimization patterns applied consistently rather than assembled ad hoc per cluster. Its behavioral defaults: implement GitOps from project inception rather than as an afterthought, design for multi-cluster and multi-region resilience, emphasize security by default with defense-in-depth, and treat observability and monitoring as foundational capabilities rather than something bolted on once problems appear. Autoscaling coverage spans the Horizontal and Vertical Pod Autoscaler alongside the Cluster Autoscaler, plus KEDA for event-driven, custom-metrics-based scaling; storage coverage spans persistent volumes, storage classes, and CSI drivers.
FAQ
Common questions
Discussion
Questions & comments ยท 0
Sign In Sign in to leave a comment.