Four sites we measured.
Not case studies. Four real runs of the same engine you get, with the site named, the findings counted by severity, and the score that came out.
The four runs, one mark per finding
Four sites, one instrument Measured
The same shape four times, so they can be compared at a glance: the score, and every finding drawn as one mark, coloured by severity. The table underneath has the same numbers written out.
Four runs · measured, not illustrative
pypi.org
50/100 · 19 findings
- 2 high
- 6 medium
- 4 low
- 7 info
Our own previous site
41/100 · 16 findings
- 1 critical
- 3 high
- 4 medium
- 4 low
- 4 info
example.com
28/100 · 39 findings
- 6 high
- 13 medium
- 13 low
- 7 info
Our deliberately broken test site
21/100 · 44 findings
- 1 critical
- 9 high
- 13 medium
- 14 low
- 7 info
The same four, written out
| Site | Score | Findings | Critical | High | Medium | Low | Info |
|---|---|---|---|---|---|---|---|
| pypi.org | 50 | 19 | 0 | 2 | 6 | 4 | 7 |
| Our own previous site | 41 | 16 | 1 | 3 | 4 | 4 | 4 |
| example.com | 28 | 39 | 0 | 6 | 13 | 13 | 7 |
| Our deliberately broken test site | 21 | 44 | 1 | 9 | 13 | 14 | 7 |
The report itself
Screenshot pending
A report, opened, with the findings worst first and the evidence attached.
Not published yet. We only publish screens of the product running, and we have not chosen which site to show here without asking its owner first. Until then this space stays empty on purpose.
What these numbers are, and what they are not
They are real. Every row is a run of this engine, and the breakdown is copied from the report’s own summary. They are checked by a test in our repository, so this page cannot quietly drift away from what was measured.
Two of the four are ours. One is the site this company used to run — it scored 41, and it had a critical. The other is a test site we broke on purpose. We show ours because a page of other people’s bad scores proves nothing about whether we are honest about our own.
They are not Melbourne tradies. These four were measured while building the engine, not gathered as customer testimonials. We are not going to relabel them as local businesses to make the page read better.
The score counts defects, not quality. It is a weighted count of what we found, so a critical finding costs as much as thirteen low ones. No number of rules turns a subtraction into a judgement.
The only run that matters to you
Is yours. The trial runs it, and every other one you want for seven days. Card on file, nothing charged until day 8, and you cancel from the dashboard at no cost.