Product quality, made inspectable
Compare runs.
Compare runs.
Calibrate judgment.
A transparent quality gate for teams that need to see where automated checks agree with human review—and where they do not.
Launch posture
Loading
Reading comparison report…
01 / Run comparisonSame cases.
Same cases.
Visible deltas.
02 / Case anatomyFind the
Find the
spread.
03 / Human calibrationAutomation is
Automation is
a hypothesis.
Use your own evidenceBring provider outputs.
Bring provider outputs.
Keep the scoring contract.
llm-eval-compare \
--dataset datasets/v1/product-support.jsonl \
--run openai=outputs/openai.jsonl \
--run anthropic=outputs/anthropic.jsonl \
--human-ratings reviews/human.jsonl \
--output reports/comparison.json