Evaluating an AI feature without relying on vibes
Building an eval set that catches regressions, where LLM-as-judge is trustworthy and where it is not, and how to run it in CI without constant flakiness.

Every post tagged Testing — practitioner notes from running Kubernetes, OpenShift and DevOps tooling in production for clients across the EU and the Gulf.
1 post
Building an eval set that catches regressions, where LLM-as-judge is trustworthy and where it is not, and how to run it in CI without constant flakiness.