EFEval field notes Reference dashboard / v0.2 View source ↗
Product quality, made inspectable

Compare runs.
Calibrate judgment.

A transparent quality gate for teams that need to see where automated checks agree with human review—and where they do not.

Launch posture Loading Reading comparison report…
01 / Run comparison

Same cases.
Visible deltas.

02 / Case anatomy

Find the
spread.

03 / Human calibration

Automation is
a hypothesis.

Use your own evidence

Bring provider outputs.
Keep the scoring contract.

llm-eval-compare \
  --dataset datasets/v1/product-support.jsonl \
  --run openai=outputs/openai.jsonl \
  --run anthropic=outputs/anthropic.jsonl \
  --human-ratings reviews/human.jsonl \
  --output reports/comparison.json