Skip to main content
An eval is a scripted benchmark against a known target. Run it to check that a command still behaves as expected, and get a written report back.

Catalog

Each run writes input.json, results.json, report.md, failures.jsonl, cost.json, and traces/ under .zenrows/evals/<run-id>/.