Design Bench
Which agent designs best?
Coding agents on real design tasks in Knighn. Every result is rendered, linted and compared with a known good answer.
Loading the results…
Real tasks
Each task starts from a design and a plain request: a landing page, an edit that must keep the look, a component with variants, a phone screen, a lint fix.
Headless runs
The agent gets only the Knighn MCP for one design file. No file tools, no web, and a time and money limit for each task.
Scored the same way
Half the score is the checks the task names, a fifth is lint, and the rest is how close the picture is to a known good answer. A run passes with every check, no lint errors and, for edits, the same look.