Kokil Thapa - Professional Web Developer in Nepal
Freelancer Web Developer in Nepal with 15+ Years of Experience

Kokil Thapa is an experienced full-stack web developer focused on building fast, secure, and scalable web applications. He helps businesses and individuals create SEO-friendly, user-focused digital platforms designed for long-term growth.

Highly Available Kubernetes Control Plane

By Kokil Thapa | Last reviewed: September 2026

A single failed control plane node should never take your cluster offline. A highly available Kubernetes control plane spreads the API server, scheduler, controller manager, and etcd across multiple machines behind one stable endpoint. That pattern matters whether you run workloads on bare metal, a VPS fleet, or a hybrid setup next to a CI/CD pipeline on Linux servers. My day job is Laravel and PHP infrastructure, but the same rules apply: eliminate single points of failure, test recovery before you need it, and monitor what actually breaks in production.

How do you architect a highly available Kubernetes control plane?

Start by naming the four components that must survive node loss. The kube-apiserver is the only entry point for kubectl, controllers, and the scheduler. etcd stores all cluster state. The kube-scheduler places pods on nodes. The kube-controller-manager reconciles desired state with reality. Lose any one of these on a single-node control plane and the cluster stops accepting changes—or stops entirely.

In a production Kubernetes architecture, you replicate all four across an odd number of control plane nodes. Odd counts matter for etcd quorum: three nodes tolerate one failure, five tolerate two. Even numbers buy you nothing extra and waste hardware.

HA Control Plane TopologyHAProxy / LBCP Node 1API ServerScheduleretcd MemberCP Node 2API ServerController Mgretcd MemberCP Node 3API ServerScheduleretcd Member
Three-node stacked etcd topology with a load balancer distributing API traffic across every control plane member

You choose between two HA topologies. Stacked etcd runs etcd as a static pod on the same node as the other control plane components. It is simpler to deploy and matches what most self-managed clusters use. External etcd places the data store on dedicated nodes. That isolates storage failures from compute but needs at least six machines—three for etcd plus three for the control plane.

For most teams without a dedicated platform group, stacked etcd on three nodes is the right default. External etcd makes sense when compliance, very large clusters, or strict I/O isolation demand it. The official kubeadm high availability guide documents both paths.

The load balancer is not optional. Worker kubelets, kubectl, GitOps agents, and CI runners all need one stable controlPlaneEndpoint. Cloud load balancers on AWS, Azure, or GCP work well. On bare metal or VPS hosts, HAProxy or Nginx in TCP mode is the usual choice. Health checks must hit /livez on port 6443—not just an open TCP socket.

TopologyMinimum NodesFailure DomainBest For
Stacked etcd3 control planeCompute + storage coupledSmall to mid-size self-managed clusters
External etcd3 etcd + 3 control planeStorage isolated from API computeLarge clusters, strict compliance
Managed control plane (EKS, GKE, AKS)0 (provider-managed)Provider SLATeams without in-house K8s ops
Stacked vs External etcdStacked (3 nodes)API + etcd on same nodeLower hardware costSimpler ops workflowNode loss hits API and etcdExternal (6 nodes)Dedicated etcd tierStronger I/O isolationMore machines to patchHigher cost and complexity
Stacked versus external etcd trade-offs when designing a highly available Kubernetes control plane

Hardware matters as much as topology. Control plane nodes need fast SSD or NVMe storage for etcd. Slow disks cause WAL fsync delays, API timeouts, and leader election storms. Budget at least 4 vCPU and 8 GB RAM per control plane node for production. Keep etcd off network-attached storage with unpredictable latency.

How do you configure HAProxy for Kubernetes API server load balancing?

Every client talks to the load balancer—not to individual API server IPs. A misconfigured LB produces intermittent 503 errors that look like application bugs. On VPS or bare-metal fleets, HAProxy in TCP mode remains a solid choice because it proxies transparently and supports deep health checks.

Install and configure HAProxy

# /etc/haproxy/haproxy.cfg
frontend k8s-api
    bind *:6443
    mode tcp
    option tcplog
    default_backend k8s-api-backend

backend k8s-api-backend
    mode tcp
    option tcp-check
    tcp-check connect port 6443
    tcp-check send GET\ /livez\ HTTP/1.1\r\nHost:\ localhost\r\nConnection:\ close\r\n\r\n
    tcp-check expect string ok
    balance roundrobin
    server cp1 10.0.1.10:6443 check inter 5s fall 3 rise 2
    server cp2 10.0.1.11:6443 check inter 5s fall 3 rise 2
    server cp3 10.0.1.12:6443 check inter 5s fall 3 rise 2

The health check is the critical detail. Raw TCP connect only proves the port listens. The /livez probe confirms the apiserver process responds. Parameters fall 3 rise 2 with inter 5s remove a dead backend within roughly 15 seconds without flapping on brief blips.

Point a DNS A record or floating VIP at the LB. That hostname becomes your permanent controlPlaneEndpoint. Changing it later means reissuing certificates and updating every kubeconfig—pain you want to avoid.

Enable the stats socket for incident debugging:

listen stats
    bind *:8404
    mode http
    stats enable
    stats uri /stats
    stats refresh 10s
    stats auth admin:change-this-password

On Ubuntu, run systemctl enable --now haproxy. Confirm all backends show UP before running kubeadm init. If you manage the underlying hosts yourself, follow the same hardening baseline you would use for any production Linux fleet—firewall rules, SSH keys, and patch cadence align with Linux system administration best practices.

How do you initialize a stacked etcd cluster with kubeadm?

kubeadm handles certificate generation, static pod manifests, and etcd bootstrapping. For HA, write a config file first. Interactive flags become error-prone once you add load balancer endpoints and multiple join commands.

Create the kubeadm configuration

# kubeadm-config.yaml
apiVersion: kubeadm.k8s.io/v1beta4
kind: ClusterConfiguration
kubernetesVersion: v1.31.0
controlPlaneEndpoint: "k8s-api.example.com:6443"
networking:
  podSubnet: "10.244.0.0/16"
  serviceSubnet: "10.96.0.0/12"
etcd:
  local:
    dataDir: /var/lib/etcd
    extraArgs:
      listen-metrics-urls: http://0.0.0.0:2381
---
apiVersion: kubeadm.k8s.io/v1beta4
kind: InitConfiguration
localAPIEndpoint:
  advertiseAddress: 10.0.1.10
  bindPort: 6443

controlPlaneEndpoint must resolve to your LB—not a single node IP. That value is baked into certificates and kubeconfig files cluster-wide. Initialize the first node with:

kubeadm init --config=kubeadm-config.yaml --upload-certs

The --upload-certs flag encrypts control plane certificates into a temporary Secret. Additional control plane nodes download and decrypt them during join. Without it, you copy PKI material manually—a process that breaks easily under pressure.

Join additional control plane nodes

Save the join command and certificate key from the init output. The key expires after two hours. Regenerate it with kubeadm init phase upload-certs --upload-certs if needed.

kubeadm join k8s-api.example.com:6443 --token abcdef.0123456789abcdef \
    --discovery-token-ca-cert-hash sha256:xxxx... \
    --control-plane --certificate-key xxxxx...

After each join, verify etcd cluster health before adding the next node:

kubectl exec -n kube-system etcd-cp1 -- \
    etcdctl --endpoints=https://127.0.0.1:2379 \
    --cacert=/etc/kubernetes/pki/etcd/ca.crt \
    --cert=/etc/kubernetes/pki/etcd/server.crt \
    --key=/etc/kubernetes/pki/etcd/server.key \
    endpoint health --cluster

All endpoints must report healthy. If a join fails halfway, run kubeadm reset on the failed node and wipe /etc/kubernetes and /var/lib/etcd before retrying. Partial state produces TLS errors that waste hours.

HA Cluster Bootstrap Sequence1. Configure LBHAProxy + /livez2. kubeadm init--upload-certs3. Join CP Nodes--control-plane4. Verify etcdendpoint health5. Install CNICalico / Cilium6. Join WorkersStandard join cmdVerify etcd after every join — partial state causes cascading TLS failuresCertificate key expires in 2 hours; re-upload if your window slips
Ordered bootstrap sequence for a highly available Kubernetes control plane with verification gates between steps

Install your CNI plugin only after all control plane nodes report Ready. Then join worker nodes with the standard join command—omit --control-plane. Tools like Kubespray automate much of this for multi-node fleets if you prefer Ansible over manual kubeadm steps.

What are the common failure modes and recovery procedures?

Knowing how a highly available Kubernetes control plane fails saves more time than memorizing a happy-path install. Three failure categories show up repeatedly in production incidents and maintenance windows.

etcd quorum loss

Three etcd members need two alive for quorum. One node down is fine. Two nodes down means the API server goes read-only or stops entirely. Monitor member count continuously:

kubectl get pods -n kube-system -l component=etcd -o wide
# Page on-call if Ready count < 2

When a node is permanently lost, remove its etcd member before rebuilding:

# List members from a healthy node
kubectl exec -n kube-system etcd-cp1 -- etcdctl member list

# Remove the failed member by ID
kubectl exec -n kube-system etcd-cp1 -- etcdctl member remove MEMBER_ID_HEX

Never reboot a long-dead etcd member and expect it to rejoin cleanly. Stale WAL data can trigger split-brain behaviour. Remove first, rebuild fresh, then join as a new member. The etcd disaster recovery documentation covers snapshot restore when quorum is fully lost.

Certificate expiration

kubeadm issues certificates with a one-year default lifetime. Expired certs produce sudden cluster-wide outages that HA cannot mask—all API servers share the same PKI.

kubeadm certs check-expiration

Renew before expiry:

kubeadm certs renew all

Restart static pods by moving manifests out of /etc/kubernetes/manifests/ and back in, or restart kubelet. Automate expiry checks in CI or monitoring. This single issue causes more surprise outages than any hardware failure I have seen on self-managed clusters.

Load balancer misrouting

When HAProxy shows backends UP but kubectl times out, suspect asymmetric routing or firewall rules. Test from every worker node:

curl -k https://k8s-api.example.com:6443/livez
# Expect "ok" from each node

Compare results across nodes. If only some fail, inspect UFW, firewalld, and SELinux policies on control plane hosts. The same systematic approach applies when debugging distributed API connectivity in Laravel apps—isolate the failing hop before changing application code.

Failure ScenarioDetectionRecoveryImpact Window
Single CP node downLB marks backend DOWNAutomatic; repair nodeNone if quorum holds
etcd leader electionAPI latency spike 2–10 sAutomaticUnder 30 seconds
Certificate expiredx509 errors from kubectlkubeadm certs renew allUntil manual fix
Two CP nodes downAPI unavailable or read-onlyRestore from etcd snapshotHours; data loss possible
LB false positiveIntermittent 503 errorsFix /livez check or routingUntil corrected
Control Plane Failure TreeSymptom Detectedkubectl timeoutx509 cert erroretcd warningsCheck HAProxy backendscerts check-expirationetcdctl endpoint healthFix LB or restart APIRenew certs + restartRemove stale memberConfirm fix: kubectl get nodes & etcdctl endpoint status --cluster
Symptom-based decision tree for rapid diagnosis of highly available Kubernetes control plane incidents

For full quorum loss, restore from a recent snapshot following the Kubernetes disaster recovery playbook. Document the procedure before you need it at 2 a.m.

How do you monitor and maintain long-term control plane health?

Building HA is a project. Keeping it HA is a discipline. The targets below mirror what I expect on any production platform—similar to the reliability bar for DevOps automation on business-critical systems.

Essential monitoring targets

  • etcd member count: Alert when ready members drop below two. Prometheus: count(etcd_server_has_leader == 1) < 2.
  • API server latency: P99 above one second on mutating requests often means etcd disk saturation.
  • Certificate expiry: Alert at 30 days. Parse kubeadm certs check-expiration in a weekly cron job.
  • LB backend state: Track HAProxy UP/DOWN transitions. Flapping backends mean network or resource problems.
  • etcd disk fsync duration: Sustained values above 10 ms mean storage is too slow. Upgrade to NVMe.

Wire these into Prometheus and Grafana. Validate alert routing with a quarterly game day—mute notifications you never act on.

Backup strategy

etcd snapshots are your last line of defence. Schedule them every four to six hours:

ETCDCTL_API=3 etcdctl snapshot save /backup/etcd-$(date +%Y%m%d-%H%M).db \
    --endpoints=https://127.0.0.1:2379 \
    --cacert=/etc/kubernetes/pki/etcd/ca.crt \
    --cert=/etc/kubernetes/pki/etcd/server.crt \
    --key=/etc/kubernetes/pki/etcd/server.key

find /backup -name "etcd-*.db" -mtime +3 -delete

Copy snapshots off-cluster immediately. A disk failure that kills etcd will also destroy co-located backups. Push to S3, GCS, or a separate backup server. Test restore quarterly—an untested backup is wishful thinking. Pair etcd snapshots with Velero for Kubernetes object backups when you need namespace-level recovery.

Upgrade procedure

Upgrade control plane nodes one at a time. Never patch all three in the same window. The sequence:

  1. Cordon the target control plane node.
  2. Upgrade kubeadm, kubelet, and kubectl packages.
  3. Run kubeadm upgrade apply on the first node; kubeadm upgrade node on the rest.
  4. Verify etcd health before touching the next node.
  5. Uncordon and proceed.

Budget 30 to 45 minutes per node. Review the CKA exam curriculum if your team needs a shared troubleshooting vocabulary before the first production upgrade.

Key Takeaways

  • Run three control plane nodes with stacked etcd and an odd member count to preserve quorum through one failure.
  • Point controlPlaneEndpoint at a load balancer that health-checks /livez, not just TCP port 6443.
  • Initialize with kubeadm init --upload-certs and verify etcd after every control plane join.
  • Automate certificate expiry checks and etcd snapshots stored off-cluster—both cause silent HA failures.
  • Remove dead etcd members before rebuilding nodes; never let stale members rejoin with old WAL data.
  • Upgrade control plane nodes sequentially and confirm etcd health between each one.

People Also Ask

How many control plane nodes do you need for high availability?

Three is the minimum for a self-managed highly available Kubernetes control plane with stacked etcd. That tolerates one node failure while keeping etcd quorum at two of three. Use five nodes only when you must survive two simultaneous failures and accept the extra hardware cost.

Can you run a HA control plane on VPS or bare metal?

Yes. HAProxy or Nginx on a separate small VM provides the API load balancer. Control plane nodes need fast local SSD, stable networking, and low latency between etcd peers. Managed Kubernetes on cloud providers hides this layer but follows the same principles internally.

What happens if two of three control plane nodes fail?

etcd loses quorum. The API server stops accepting writes and may become fully unavailable. Recovery requires restoring from a recent etcd snapshot or rebuilding the cluster. This is why off-cluster backups and tested restore runbooks are non-negotiable.

Is stacked etcd safe for production?

Stacked etcd is production-grade for most clusters when nodes use NVMe storage, you maintain three or five members, and you monitor fsync latency. External etcd adds isolation for very large or compliance-heavy environments at the cost of double the machine count.

Plan your production control plane with confidence

A highly available Kubernetes control plane is not magic—it is three nodes, one stable LB endpoint, disciplined certificate and backup hygiene, and runbooks you have actually tested. Skip any of those and your HA setup becomes a single point of failure with extra steps. If you are designing a production cluster, hardening an existing control plane, or integrating Kubernetes into a broader platform alongside applications like those in our Laravel and Livewire booking portfolio, validate your YAML configs with our JSON and YAML formatter before apply. For hands-on help with cluster design, support and maintenance, or infrastructure hardening, contact us to discuss your requirements. You can also reach out directly with specific architecture questions.

Frequently Asked Questions

A configuration with three or more etcd members and multiple API servers behind a load balancer, ensuring cluster management survives individual node failures without downtime.

Minimum three nodes to maintain etcd quorum during one failure; five nodes tolerate two simultaneous failures but increase operational cost and complexity significantly.

Etcd uses Raft consensus requiring (N/2)+1 votes; even numbers provide no extra fault tolerance over the next lower odd number while increasing split-brain risk.

In my experience managing infrastructure, stacked etcd works reliably for clusters under 50 nodes and simplifies operations. External etcd adds deployment complexity but isolates API server load from consensus traffic, making it necessary only when control plane resources face contention or you manage multiple clusters sharing etcd backing stores. For most Nepal-based deployments on limited budgets, stacked reduces both cost and failure domains.

Use kube-vip or keepalived for bare metal VIPs, or cloud provider LBs for managed environments. The load balancer must forward TCP 6443 to all healthy API servers and perform active health checks against /livez endpoints. Avoid DNS-only round-robin as clients cache stale IPs during failover. On Ubuntu servers I manage, kube-vip in ARP mode provides sub-second failover without external dependencies, costing zero additional NPR compared to managed alternatives.

Remaining nodes maintain etcd quorum and continue serving API requests through the load balancer. Automatic leader election occurs within seconds. Workloads remain unaffected as worker nodes communicate via the VIP. However, certificate renewal, scheduled jobs, and controller reconciliation pause briefly. Monitor etcd member health immediately; running degraded with two of three nodes means losing another causes total cluster failure. Always replace failed nodes before performing maintenance elsewhere.

Three minimal VMs (2 vCPU, 4GB RAM) cost roughly Rs 15,000–25,000/month (~USD 110–185) on typical cloud providers. Bare metal requires higher upfront investment but eliminates recurring compute costs. Budget an additional 20% for load balancer, storage, and monitoring overhead.

Yes, using kubeadm upgrade workflow. Add new control plane nodes sequentially with kubeadm join --control-plane, verify etcd membership and API server health after each addition, then update your load balancer backend pool. Never remove the original node until at least two new members are healthy and synced. Back up etcd snapshots before starting. This process takes 30–60 minutes per node depending on network and storage performance.

Disk latency exceeding 10ms causes leader elections and API timeouts. Network partitions between members break quorum. Clock skew over 500ms corrupts consensus. Insufficient disk space triggers compaction failures. In production systems I have debugged, slow NVMe degradation was the silent killer—monitoring showed healthy CPU/memory while etcd fsync durations climbed gradually. Always use dedicated SSDs, enable etcd metrics, and alert on WAL fsync p99 latency exceeding 5ms.

Use etcdctl snapshot save on a single member during low-traffic periods; never backup all members simultaneously. Store encrypted snapshots off-cluster. Restoration requires stopping all API servers, restoring to one member with etcdctl snapshot restore, then rejoining other members as fresh learners. Test restores quarterly in staging. On legal-tech portals handling sensitive documents, I automate daily snapshots to S3-compatible storage with versioning, ensuring recovery point objectives stay under one hour regardless of failure type.

Each API server needs certs signed by the same CA with SANs covering the VIP, individual node IPs, and localhost. Etcd members require separate peer and client certs. Kubelet and controller-manager need distinct identities. Certificate rotation must happen before expiry across all nodes simultaneously or rolling with load balancer draining. Use cert-manager or kubeadm's built-in rotation. Expired certs cause silent API failures that mimic network issues; always validate cert validity during troubleshooting.

Track etcd member count, leader changes, DB size, and WAL fsync latency via Prometheus. Alert on API server request duration p99 exceeding 1s, failed authentication attempts, and node NotReady status. Use Grafana dashboards specifically designed for HA control planes showing quorum status visually. Synthetic probes hitting the VIP endpoint catch load balancer misconfigurations that individual node metrics miss. Logging aggregation helps correlate events across nodes during incidents. Without proper observability, HA becomes false confidence rather than genuine resilience.

Only if downtime directly loses revenue or violates compliance. Single-node clusters with regular backups recover in minutes for non-critical workloads. HA adds significant operational burden: certificate management, upgrade coordination, troubleshooting complexity. For most Nepal SMB sites I build, managed Kubernetes or simpler orchestration suffices. Reserve HA for platforms where four-hour recovery windows are unacceptable, like payment processing or legal service portals with strict SLAs. Calculate actual business impact before investing in HA infrastructure.

Etcd peers need reliable low-latency connectivity on ports 2379 and 2380. API servers communicate on 6443. All nodes require stable DNS resolution and synchronized time via NTP. Cross-datacenter deployments introduce latency that breaks Raft timing assumptions; keep all control plane nodes within the same region or availability zone. Firewall rules must allow bidirectional traffic between members. Network jitter causes more HA failures than hardware faults in my experience; use dedicated VLANs or security groups isolating control plane traffic from application workloads.

Upgrade one node at a time: drain, cordon, upgrade kubeadm/kubelet/kubectl, uncordon, verify health, repeat. Etcd upgrades follow separate compatibility matrices. Never upgrade multiple nodes simultaneously. Pre-upgrade validation with kubeadm upgrade plan catches breaking changes. Maintain rollback capability by keeping previous binaries accessible. Schedule upgrades during maintenance windows despite HA claims; cascading failures during upgrades cause extended outages. Document exact versions tested together. On production clusters, I stage upgrades in identical test environments first, validating workload compatibility before touching live systems.

Share this article

0 Comments

Leave a comment

Your email is not published. Comments appear once they have been read. Sign in to have your details filled in.

Quick Contact Options
Choose how you want to connect me: