LLM-as-judge

Appears in 3 tutorials

Using another LLM to score an output against criteria (also used in evaluation).

As used in AI Agents →

Using another LLM to score an output against criteria (also used in evaluation).

As used in Prompt Engineering →

Using another LLM to score or compare outputs against criteria — a fast, scalable way to run evals.

As used in AI Production Engineering →

Using another LLM call to grade an output against a rubric. Scales nuanced evaluation; must be validated against humans. (Mod 4)