Evaluating an AI feature without relying on vibes
Building an eval set that catches regressions, where LLM-as-judge is trustworthy and where it is not, and how to run it in CI without constant flakiness.

Every post tagged LLM — practitioner notes from running Kubernetes, OpenShift and DevOps tooling in production for clients across the EU and the Gulf.
2 posts
Building an eval set that catches regressions, where LLM-as-judge is trustworthy and where it is not, and how to run it in CI without constant flakiness.
Prompt caching, model tiering, batching and effort control: four levers that cut spend on an LLM feature, in the order that returns most for least risk.