200 OK is not health

Route /health, returns ok, everyone happy. That is the first version of every health endpoint on the project we are moving to Kubernetes. Then somebody pastes the same path into livenessProbe and readinessProbe, and a thirty second database hiccup becomes a long evening. The two probes ask different questions. Readiness asks: should this pod get traffic right now. Here it is correct to check dependencies. Database unreachable, cache cold, migrations still running: answer no. Kubernetes takes the pod out of the Service, traffic goes to the others, the pod returns when the world improves. Failing readiness is cheap and reversible. A polite “not now”. ...

August 7, 2019 · 2 min · Murat Useinov

What Redis does when memory ends

Users logged out at random. Not all, not always. Only sometimes, only in the afternoon. Afternoon is when traffic peaks. Traffic peaks fill the cache. The cache lived in the same Redis as the sessions, the instance hit maxmemory, and the eviction policy was allkeys-lru. Redis did exactly what we asked: threw away the least recently used keys, and some of them were sessions of people who went for lunch. ...

April 2, 2019 · 2 min · Murat Useinov