200 OK is not health

Route /health, returns ok, everyone happy. That is the first version of every health endpoint on the project we are moving to Kubernetes. Then somebody pastes the same path into livenessProbe and readinessProbe, and a thirty second database hiccup becomes a long evening. The two probes ask different questions. Readiness asks: should this pod get traffic right now. Here it is correct to check dependencies. Database unreachable, cache cold, migrations still running: answer no. Kubernetes takes the pod out of the Service, traffic goes to the others, the pod returns when the world improves. Failing readiness is cheap and reversible. A polite “not now”. ...

August 7, 2019 · 2 min · Murat Useinov