Loupe
Loupe measures a rendered web page and tells a coding agent what is out of range, and by how much. It is a research preview by CodePawl.
A coding agent writes CSS and does not see the result. Loupe renders the page, measures it, and returns lines like this one:
font size relative to 16 px of "Electronics": too low. Measured 13 / 16 = 0.812;
in range is 1.31 to 1.44, so raise the ratio by about 63%
The agent reads the lines and edits the code. Loupe then measures the page again.
What is in this repository
| File | What it is |
|---|---|
loupe.py |
The command line tool. It renders a page at one or more window widths and prints what to fix. |
collect_page.cjs |
The renderer. It uses Playwright to find text blocks, buttons, pictures and repeated groups. |
page_audit.py, group_checks.py, rule_checks.py |
The checks. Each check is one ratio and a range. |
read_runtime.py, read_v2/read.safetensors |
The trained read: about 121 thousand parameters. |
block_measure.py, pixel_measurements.py, block_maps.json |
Pixel reading of contrast, font size and line height. |
benchmark.py, trusted_checks.json |
The benchmark and the list of checks that pass it. |
sensitivity.py, timing.py |
The fault size test and the speed test. |
charts_web/ |
The chart page and the script that renders it. |
charts/ |
The charts on this page. |
reports/ |
The numbers behind every result on this page. |
The training code and the benchmark pages are not in this repository. The benchmark pages are screenshots of other people's sites.
Use it
Install the parts:
pip install torch numpy pillow safetensors
npm install playwright
npx playwright install chromium
Audit a page:
python loupe.py audit https://example.com --widths 1440,375 --advice
--widths sets the window widths. --advice also prints the checks that have not passed the benchmark.
To use an installed Chrome, set LOUPE_CHROME to its path.
To hold a page to an earlier good version of itself, collect the good version and pass it as the reference:
import page_audit
report = page_audit.audit(page_directory, reference=good_version_directory)
print(page_audit.flagged_lines(report))
How it works
- Playwright renders the page. The renderer records where each text block, button, picture and repeated group is.
- Loupe takes font size and line height from the browser. It reads contrast from the screenshot pixels.
- Each check becomes two numbers and a range, for example heading size, text size, and 1.3 to 4.0.
- The trained read answers "in range", "too low" or "too high" with a probability.
- Loupe prints one line for each check that is out of range. The line has the measured numbers and the change to make.
The read is exact near a bound. In its training gate it made no error further than 0.3 percent from a bound.
Results
All numbers are measured. The files in reports/ hold them.
Benchmark
The benchmark has 33 pages from 11 public sites. Loupe was not tuned on these pages. Each page gets 22 kinds of injected fault, one at a time. An example is one card moved down by 22 px.
A check is trusted when it passes two gates. It must catch at least 90 percent of its faults. It must flag at most 10 percent of what it checks on untouched pages.
Seven checks pass:
| Check | Faults caught | Flagged on untouched pages |
|---|---|---|
| Row alignment | 22 of 24 | 0.9% |
| Row gap evenness | 24 of 24 | 6.0% |
| Text contrast | 28 of 30 | 3.3% |
| Sideways overflow | 19 of 20 | 3.0% |
| Body text size | 25 of 26 | 9.3% |
| Text touching the page edge | 31 of 32 | 0.2% |
| Main heading in the first screen | 24 of 26 | 0.0% |
Fifteen checks do not pass yet. Loupe still measures them and marks them as advice. Four of them are weak in both ways: text overlap, heading size consistency, heading proximity and clipped text. The audit no longer runs the first three of those.
Against a vision model
We cut a 1440 by 900 window around each fault, from the faulty page and from the untouched page. We asked Claude Sonnet which of the 22 issues it saw in each picture. Loupe answered the same question from its audit. Each fault kind has eight windows.
| Said yes on 176 faulty windows | Said yes on 172 untouched windows | |
|---|---|---|
| Claude Sonnet, picture only | 109 | 10 |
| Loupe | 154 | 22 |
Loupe caught more faults. It also said yes more often on untouched windows. A yes on an untouched window is not always wrong. Real pages have small buttons and light headings.
Loupe was ahead on uneven card heights (8 of 8 against 1 of 8). It was also ahead on a main heading pushed below the first screen (8 of 8 against 1 of 8). Sonnet was ahead on tight line spacing (8 of 8 against 6 of 8) and on overlapping text (7 of 8 against 5 of 8).
Sonnet got the picture only. Loupe got the rendered page.
An earlier run with four windows for each fault kind is in reports/versus_sonnet_first_4_per_fault.json.
Speed and size
The read has 121,153 parameters. The weights file is 474 KiB. On 12 live pages, Loupe judged a page in a median of 0.95 seconds on a CPU. A page had a median of 836 checks.
Loading and rendering the live page took a median of 8.0 seconds. That time includes the network. Sonnet took a median of 9.3 seconds for one window, without loading or capturing the page.
How small a fault it catches
We injected three geometric faults at 2, 4, 8 and 16 px.
| Fault | 2 px | 4 px | 8 px | 16 px |
|---|---|---|---|---|
| One card moved sideways | 11 of 24 | 23 of 23 | 24 of 24 | 24 of 24 |
| One card moved down | 0 of 24 | 6 of 24 | 16 of 24 | 22 of 24 |
| One heading indented | 0 of 17 | 0 of 17 | 0 of 17 | 15 of 17 |
A heading indent counts from half a font size, by design.
On a human-judged benchmark
This is the one test here that we did not design. It is the pairwise human evaluation from Design2Code. Each row has a reference page and two pages that models wrote to copy it. Five human judges voted on which copy is closer. We fixed Loupe's rule before we looked at any result. A copy scores the share of the reference's measurements that it matches within 5 percent.
| Agrees with the judges on 582 rows | |
|---|---|
| Loupe | 291 (50.0%) |
| Mean pixel difference between the screenshots | 346 (59.5%) |
Loupe did not beat the pixel baseline. It called 120 rows a tie, and each of those counts as wrong. On the 462 rows where it chose a side, it agreed with the judges on 63.0%.
The rule was too strict for this task. A copy that people find close still differs from the reference by more than 5 percent in most measurements.
We replaced the rule with distances. Loupe matches each text block of the reference to the block with the same text in the copy. It then measures how many blocks are missing, and how far the matched blocks are in font size, weight, colour, width and position. A small logistic model turns the two copies' distances into a choice. We fit it on the even rows and tested it one time on the odd rows.
| On 288 held-out pairs | Agrees with the judges |
|---|---|
| Every element within 5 percent | 52.1% |
| Mean pixel difference | 63.9% |
| Loupe distances | 76.4% |
With 288 pairs, each number is good to about 5 points either way.
The code is similarity.py.
We also gave Claude Sonnet the same question on 100 of the held-out pairs, with three screenshots for each pair. Sonnet agreed with the judges on 83%, Loupe on 73% and the pixel difference on 65%. Sonnet took a median of 10.6 seconds for each pair. On this question the vision model is ahead of Loupe.
The reference mode now reports the three things that weighed most for the judges: missing text, a changed colour and a moved left edge.
Contrast
We rendered each page two times, with its text and without it. The difference shows the text pixels and the real background. On 3,737 text blocks from 33 unseen pages, Loupe flagged 0.1 percent of the blocks that were in range. It missed 17 of 123 blocks that had low contrast.
One repair run
We took a page that passed, and changed eleven CSS values by hand. Loupe held the page to its earlier version, within 5 percent for each element.
Claude Sonnet did not see the page. It got the lines from Loupe and nothing else. The page went from 335 of 410 checks in range to 399 of 399 in two rounds. The video shows this run. This is one page and one run.
Limits
- Loupe needs to render the page. It cannot audit a screenshot that it did not make.
- The default ranges come from written design rules. Many real sites do not follow those rules, so most checks are advice only.
- The reference mode works for one page at two points in time. It does not work for "make this page look like that site".
- The pass or fail reference rule does not match human judgement of similarity. The distances in
similarity.pydo better, and they are not part of the audit yet. - Font size read from pixels is not reliable on unseen pages. Line spacing flagged 17 percent of in-range text and missed 39 percent of faults. Loupe uses the browser values for that reason.
- Three earlier repair runs ended with every check in range while the page still had a visible fault. Each time one measurement was missing.
- In 8 repair runs we changed 48 CSS values at random. The agent put 25 back and set 7 to other values. 16 stayed broken. Three of those had no visible effect. Most of the others were colours and the line height of one-line text. Loupe reported every check in range in 7 of the 8 runs. Its score does not prove that a page is restored.
- Loupe does not measure colour harmony, visual hierarchy, or hover and focus states.
- Loupe does not measure forms, user flows or the quality of the words.
- Loupe checks a page against a range. It does not know if a page is beautiful.
License
Apache 2.0. The demo page in the video uses photos from Unsplash.




