CodeRush.run home

For coding agents · runs on your machine

A screenshot shows it’s wrong.
CodeRush shows why.

Show a vision model a screenshot and it can tell you a page looks off. It cannot see an uncaught error or a request that failed, tell a dark scene from a canvas that never drew, or know which line to change. One CodeRush check renders the page in a real browser and returns the screenshots together with those facts, each finding tied to the source line and the CSS rule behind it. The model judges what looks wrong; the report says why, and where.

  • Screenshots and measured facts
  • Every finding has its line
  • MCP, command line or HTTP

Eyes and instruments

Each catches what the other cannot.

A vision model is good at what people are good at: noticing a cramped layout, an overlap, a heading that looks wrong. It is poor at facts that are not in the pixels, and it cannot know which line of code produced what it sees. Measurement is the opposite. So CodeRush does not replace the model. It hands the model what the model is missing.

What each catches on the same page
ProblemA screenshot sent to a vision modelA CodeRush check
Cramped, crowded, misaligned or uglySees itCannot judge taste; the brief asks the model, by selector
A black frame: a night scene, or a canvas that never drew?Guessesflat_color and blank, read from the screenshot’s own pixels
A thrown error, a console error, a failed or blocked requestInvisible in the pictureerrors, console, network.failed, network.blocked
A button that does nothing when clickedA still picture cannot show itinteract clicks it and reports what changed
Content wider than the screenSometimes, never by how muchsticking_out: the element, how far it reaches, its line and the rule that set its width
Text cut off, or too faint to readSometimesclipped_text with the box it needs; unreadable_text with both colours and the ratio
Buttons too small to tapA rough guesssmall_targets under 24×24px, with the exemptions WCAG allows
Which line to changeCannot knowEach finding carries its source line and the CSS rules behind it, with their lines

So send both. Every report carries a brief: the measured facts first, marked as facts; then where each element sits on the screenshot, as a map of selectors, boxes and lines; then a request for only what measurement cannot see. The model’s answer stops being “the cards look crushed” and becomes a selector, which the map turns into the line to edit. Over MCP the agent reading the report is usually that vision model already: it gets the screenshots as images and the brief as text in the same answer.

One check, the whole picture

Read the hint first.

Send source (HTML, JSX, TSX, JavaScript, Markdown, SVG, CSS, JSON) or the URL of a dev server. CodeRush renders it at the widths and themes you ask for and returns one report. It lists uncaught errors, the console, failed and blocked requests, and the headings and controls on the page, each with a ref such as @e2. It flags sideways overflow and names the element that causes it, text you can barely see with both colours, text cut off by its box, tap targets under 24×24px, broken images, and blank or one-colour screenshots. For HTML source every finding also names the line the element was written on and the CSS rules behind it. At the top, one hint sentence says what to fix first; the brief is the same report written for a vision model.

To test behaviour, interact replays the page with steps such as {"click":"@e2"} and says what changed: how much of the screenshot moved, where, and which controls appeared or went.

Three ways in, one reportshell
# MCP, for Claude Code (from a checkout, after npm install)
claude mcp add coderush -- node /path/to/coderush-run/tools/mcp.mjs

# command line: JSON when piped, exit 1 when it finds problems
node tools/coderush.mjs check index.html --viewports 390,1280 --themes light,dark

# local HTTP service
npm run serve
curl -s -X POST http://127.0.0.1:8765/check \
  -H "Authorization: Bearer $(node tools/coderush.mjs token)" \
  -H 'Content-Type: application/json' \
  -d '{"source":"<h1>hello</h1>","viewports":[390,1280]}'

The npm package is not published yet, so today it runs from a checkout of the repository. Read the full contract at /llms.txt

Measured, not promised

Thirty-four planted bugs, six clean pages.

The repository carries a benchmark of 34 pages with one planted UI bug each: a thrown error, a console error, a missing file, a blocked request, a phone-width overflow, dark-mode text that disappears, a blank page, a canvas that never draws, a dead button, a broken image, a JSX syntax error, a runaway loop, a label cut off by its box, a heading that loses its second line on a phone, a 16px dismiss button and 18px pagination links. It also has 6 pages written to be clean, one of them built from the things that must not be flagged: an intentional ellipsis, a link inside a sentence, labelled fields and a browser-sized button. Each page gets one check at 390 and 1280 pixels, light and dark. On 12 September 2026 the report caught all 34 bugs, and the hint named 33 of them directly. The 34th named the blocked request that broke the image. All 6 clean pages came back with no problems, and across the site’s 32 demos it flagged 3 pages, each a real bug.

It is our own benchmark, and it measures the report, not an agent. The first run caught 27 of the 30. The misses led to three fixes: a JSX file that did not compile, and a component that never mounted, had both passed as clean pages, and a body hidden with display:none was reported as a flat colour instead of blank.

Why it stays on your machine

Untrusted source, bounded blast radius.

Generated code is untrusted code. Every check runs in a fresh browser process, and the parent kills the whole process group when its deadline passes. By default the page may load from public code CDNs and nothing else; every blocked request is listed in the report. Checks live in memory and vanish when the process exits. Nothing is shared unless the agent asks for a link, and a failed render never gets one.

The HTTP service binds to 127.0.0.1 and answers only when the Host header names this machine, which stops a web page that points its own domain at your computer. POST requests need a token the service writes to a file only you can read. When a checked file loads its folder, hidden files such as .env and anything outside the folder are refused.

What it does not stop: DNS lookups from the page, and anything the page sends to a host you allowed.

Written for a reader that is not a person

Constraints before code.

/llms.txt is the agent-facing contract: the tools, the inputs, the report, the limits and the boundaries, in the order a model needs them. The repository also ships a skill file that teaches the loop in one page, and the status tool tells an agent the formats, limits and allowed CDNs before it writes a line.

If your agent has a browser instead of a shell, the hosted pad takes a Rush link — source in the fragment, rendered on open — which is a convenient way to hand a human the exact thing the agent built. And the extension covers the case where the code is sitting in a page you are already reading.

Questions

Why not just send a screenshot to a vision model?
A screenshot answers one question: does it look right? It cannot show an uncaught error, a request that failed or a canvas that never drew, it shows overflow and cut-off text only sometimes and never by how much, and the model cannot know which line produced what it sees. Send both. A check returns the screenshots with those facts, each tied to its source line and CSS rule, and a brief that asks the model only for what measurement cannot see.
Does CodeRush call an AI model?
No. It renders, measures and reports. Over MCP the agent reading the report is usually a vision model already: it gets the screenshots as images and the brief as text in the same answer. From the command line or HTTP you send them to whichever model you use.
Which findings get a line number?
Every layout finding and every entry in the map, when the source is HTML: the line the element was written on, plus the CSS rules behind the problem with their lines. JSX, Markdown and the other formats are turned into pages whose lines are not yours, so their findings say line null instead of a wrong one.
Does it need an account or an API key?
No. The MCP server and the command line need nothing. The local HTTP service asks for a token it writes to your home folder when it starts, so other web pages and other users on the machine cannot drive it.
Which agents can use it?
Any agent that speaks MCP or can run a command: Claude Code, Codex, Cursor and others. The report is the same JSON on all three surfaces.
What does one check report?
Uncaught errors, console errors and warnings, failed and blocked requests, the headings and controls on the page with refs to click, sideways overflow with the element that causes it, text with contrast under 3:1, text cut off by its box, tap targets under 24×24px, broken images, blank or one-colour screenshots, a map of where elements sit on each screenshot, the source line and CSS rules behind each finding, a brief for a vision model, and the screenshots themselves, per viewport and theme.
What happens if generated code hangs?
Every check runs in a fresh browser process with a hard deadline, and inline loops stop after three seconds of looping. The report says what ran out of time.
Is anything kept?
Checks live in the memory of the process that ran them and vanish when it exits. Nothing is shared unless the agent asks for a link. There is no hosted version of the service.

Eyes and instruments.

Give the model the picture, and the facts it cannot see.