Skip to content
← All posts

evaluation

Deciding whether an LLM output is any good — rubrics, judges, offline datasets, and the gap between scoring well on a benchmark and working for your users.

14 posts

Other subjects