How would you use Kubernetes Horizontal Pod Autoscaler (HPA) to auto-scale pods based on incoming load?
💡 Model Answer
Kubernetes HPA auto‑scales pods by monitoring metrics. To scale based on incoming load, you can use CPU/memory metrics or a custom metric such as request rate or queue depth. First, enable the metrics‑server or install a Prometheus Adapter to expose custom metrics. Then create an HPA resource specifying minReplicas, maxReplicas, and a targetCPUUtilizationPercentage or targetValue for the custom metric. The HPA controller polls the metric every 15 seconds and adjusts the replica count. You can also set a cooldown period to avoid rapid oscillations. For queue‑based scaling, expose the queue depth as a custom metric via Prometheus and configure HPA to target a desired depth. This ensures pods scale up when load spikes and scale down when idle, keeping resource usage efficient.
This answer was generated by AI for study purposes. Use it as a starting point — personalize it with your own experience.
🎤 Get questions like this answered in real-time
Assisting AI listens to your interview, captures questions live, and gives you instant AI-powered answers on a discreet on-screen overlay.
Get Assisting AI — Starts at ₹500