You should have an agent that grades your agents
Try Warp Factories: warp.dev/factories/request-access
Let's measure the quality of coding agents (Claude Code, Codex, etc) with scoring agents.
These use LLM-as-a-judge to measure output quality along metrics you care about, like code quality, efficiency, verbosity, and task compliance.
These feed self-improvement and benchmarking workflows to improve your software factory automatically.
Warp
Warp is the terminal with AI and your dev team's knowledge built-in. ...