Design Bench

Which agent designs best?

Coding agents on real design tasks in Knighn. Every result is rendered, linted and compared with a known good answer.

Loading the results…

Real tasks

Each task starts from a design and a plain request: a landing page, an edit that must keep the look, a component with variants, a phone screen, a lint fix.

Headless runs

The agent gets only the Knighn MCP for one design file. No file tools, no web, and a time and money limit for each task.

Scored the same way

Half the score is the checks the task names, a fifth is lint, and the rest is how close the picture is to a known good answer. A run passes with every check, no lint errors and, for edits, the same look.