January 29, 2026•By Andrés González•7 min read

    10 Practical Tips for a Scalable, Stable, and Cost-Efficient Kubernetes Cluster

    Lessons from running Open edX and Owly in production: databases, StatefulSets, availability zones, requests and limits, Karpenter, HPA, PDBs and probes.

    10 Practical Tips for a Scalable, Stable, and Cost-Efficient Kubernetes Cluster

    This article is written for engineers, architects, and technical leaders running Kubernetes in production — especially in SaaS and multi-tenant environments.

    After many years deploying Open edX, Owly (our AI assistant), and other production workloads in the cloud, I’ve accumulated a long list of lessons learned about how to run Kubernetes clusters that are performant, resilient, and affordable.

    These are practical recommendations, based on real incidents and trade-offs. Each topic could easily deserve a full article on its own. Not every tip applies to every use case, but most are relevant for typical web-based SaaS applications, multi-tenant platforms, and business workloads.


    1. Keep databases out of the cluster

    I won’t say that you can’t run databases inside Kubernetes—although I’m tempted to.
    Unless your product is specifically about running databases in Kubernetes, I’ll just say: if you want to sleep at night, keep your databases outside the cluster.

    For OLTP workloads, managed database services from your cloud provider are usually the best choice:

    • Infrastructure is tuned for database performance

    • Built-in redundancy, backups, and maintenance

    • You can shut down or recreate your cluster without losing data

    • Database load won’t compete with application workloads

    If managed services are not an option, run databases on dedicated infrastructure, not inside the same cluster that hosts your application pods.


    2. Be extremely careful with StatefulSets

    Kubernetes scalability is built around the idea that workloads are stateless and disposable. Pods can be created, destroyed, and rescheduled anywhere at any time.

    Stateful workloads break that model.

    Using PersistentVolumes introduces several challenges:

    • You must handle concurrent access and locking yourself

    • Volumes are usually tied to a single availability zone

    • Kubernetes does not natively handle replication or backups of PersistentVolumes; this depends on the underlying storage provider

    Whenever possible, prefer:

    • Object storage (S3, GCS, Azure Blob)

    • S3-compatible systems like MinIO

    • Ephemeral volumes, if data loss is acceptable

    Real incident:
    We once had a pod running Caddy storing TLS certificates on a PV. When the pod needed to restart, the only available nodes were in a different AZ. The volume affinity prevented scheduling, and the entire application went down.

    The fix: move certificate handling to cert-manager + NGINX ingress and remove the PV entirely, while keeping Caddy just as a proxy. The pod can now reschedule anywhere.


    3. Span your cluster across multiple availability zones — and size your networks properly

    A single availability zone is usually reliable—but not infallible.

    Spanning your cluster across two or more AZs provides a good balance between resilience and cost, without the complexity of multi-region deployments.

    Two important points to remember:

    • Pods and nodes consume IP addresses

    • System and monitoring pods count too

    If your CIDR range is too small, you will run out of IPs sooner than expected.

    Rule of thumb:
    Size your VPC and subnets for your maximum expected number of nodes and pods, not your initial cluster size. If you run out of IP addresses, you’ll need to add more networks or subnets — which is much easier to do early than after the cluster is live.


    4. Set up monitoring early

    Once your cluster grows, tuning it purely from the CLI becomes painful.

    At minimum, install:

    • metrics-server (required for HPA and autoscaling)

    A solid monitoring stack:

    • Prometheus for metrics

    • Grafana for visualization

    • Loki for logs

    You can’t optimize what you can’t see.


    5. Always define CPU and memory requests and limits

    Kubernetes scheduling decisions are based on requests, not actual usage.

    Limits, on the other hand, define what happens when things go wrong.

    Example:

    resources:
      requests:
        cpu: "500m"
        memory: "2Gi"
      limits:
        cpu: "1000m"
        memory: "4Gi"
    

    The sum of the requests of all pods allocated to a node will be less than or equal to the node’s capacity, both for CPU and memory.

    Pods can consume resources above their requests up to the configured limits, competing with other pods for the node’s available resources.

    In this example:

    • The pod can use more than 500m CPU, but will be throttled at 1000m

    • CPU throttling makes applications slow and prone to timeouts

    • Memory behaves differently: if the pod exceeds 4Gi by even a single byte, it will be immediately killed (OOM)

    Key points:

    • Requests determine where pods can be scheduled

    • CPU limits cause throttling

    • Memory limits cause OOM kills

    • Without memory limits, one pod can destabilize an entire node

    As guardrails, define:

    • ResourceQuota per namespace

    • LimitRange defaults for containers

    You can also enable VPA in recommendation (“off”) mode to observe real usage patterns before enforcing changes.


    6. Use a private container registry

    Public registries work—until they don’t.

    We once rotated multiple nodes at the same time. New pods failed to start because Docker Hub rate limits were exceeded.

    • Anonymous pulls: ~100 per 6 hours

    • Authenticated pulls: ~200 per 6 hours

    In real clusters, that limit is surprisingly easy to hit.

    If you’re on AWS, ECR is the obvious choice. Otherwise, use any private registry or a paid Docker Hub plan.


    7. Use Karpenter — but isolate system workloads

    Node provisioning matters.

    Traditional managed node groups scale based on node-level metrics.
    Karpenter, instead, provisions nodes based on pod resource requests, which gives much finer control.

    However, Karpenter runs inside the cluster it manages.

    That’s dangerous.

    Best practice:

    • One small, managed node group for system pods

    • One or more Karpenter-managed node groups for applications

    System nodes should host:

    • Karpenter

    • CoreDNS

    • Metrics server

    • Other critical components

    Use taints and tolerations to enforce separation.

    Also: avoid burstable (t*) instances for nodes. CPU credits can run out at the worst possible moment.


    8. Configure HPA carefully (especially for memory)

    HPA scales pods based on requests, not limits.

    You can scale on:

    • CPU (recommended)

    • Memory (use with caution)

    A CPU burst can easily go from almost zero to the limit configured (or hit 100% of the node’s CPU if not limited) in a fraction of a second and stay there for a long time. This can trigger multiple replicas if not controlled properly. Even a simple bug (for example, an unintended endless loop) can bring the whole cluster down.

    Memory-based HPA is tricky because baseline memory usage is rarely zero. A pod at rest can stay at 0 CPU but will usually consume some amount of memory. If the threshold is too close to that baseline, pods will scale up but never scale down. That baseline memory usage is then multiplied by the number of replicas, consuming node resources unnecessarily.

    Rule of thumb:
    Set memory targets well above the baseline usage—often more than 2×—to allow scale-down behavior.

    Takeaway: CPU-based HPA is usually safe; memory-based HPA requires deep knowledge of baseline usage and should be used sparingly.


    9. Protect availability with PodDisruptionBudgets

    PodDisruptionBudgets (PDBs) control how pods are evicted during:

    • Node drains

    • Karpenter consolidation

    • Rolling updates

    Without PDBs, a single-replica deployment can disappear temporarily during rescheduling.

    Example:

    apiVersion: policy/v1
    kind: PodDisruptionBudget
    metadata:
      name: myapp-pdb
      namespace: mynamespace
    spec:
      minAvailable: 1
      selector:
        matchLabels:
          app.kubernetes.io/name: myapp
    

    Be careful: overly strict PDBs can block deployments and autoscaling. If things get stuck, scale replicas up temporarily to restore flexibility.


    10. Use liveness and readiness probes correctly

    Probes tell Kubernetes whether a pod:

    • Is alive

    • Is ready to receive traffic

    Best practices:

    • Liveness: cheap and fast (process check, simple endpoint)

    • Readiness: application-level health (dependencies, DB connectivity)

    • For slow startups, consider startupProbe

    Whenever possible, expose:

    • /live → returns 200 OK immediately

    • /ready → validates dependencies

    Example:

    livenessProbe:
      exec:
        command:
          - sh
          - -c
          - "pgrep -f uwsgi"
      initialDelaySeconds: 10
      periodSeconds: 10
      timeoutSeconds: 3
      failureThreshold: 3
    readinessProbe:
      httpGet:
        path: /health/
        port: 5000
        httpHeaders:
          - name: Host
            value: backend.myapp.com
          - name: X-Forwarded-Proto
            value: https
      initialDelaySeconds: 30
      periodSeconds: 30
      timeoutSeconds: 10
      failureThreshold: 3
    

    Here, pgrep -f uwsgi is executed to check if the uWSGI service is running. Then the /health/ endpoint is requested on port 5000 to verify that the application is responding. This endpoint should return 200 OK.

    This dramatically improves resilience and rollout behavior.


    Final thoughts

    Running Kubernetes in production is not trivial. But once you understand how scheduling, scaling, and failure modes really work, it becomes an incredibly powerful platform.

    Most outages are not caused by Kubernetes itself—but by implicit assumptions about how it behaves.

    Kubernetes is opinionated. Production success comes from understanding those opinions — and designing with them, not against them.

    KubernetesOpen edXDevOpsKarpenterCloud costs
    A

    Andrés González

    CTO & Cofounder at Aulasneo