Real runs, real numbers.
The scores and tables below are read straight from the committed eval run. They update themselves when the engine is re-measured, so the page can never quietly disagree with the numbers it ships.
English: a productivity blog post, score 16 to 77
| check | before | after | |
|---|---|---|---|
| sentence-length stdev | 3.3 | 1.3 | still failing |
| em dashes | 1 | 1 | still failing |
| AI-inflation vocabulary | 12 | 0 | fixed |
| stock AI phrases | 2 | 0 | fixed |
| negative parallelism | 1 | 0 | fixed |
| participle tails | 3 | 0 | fixed |
| significance inflation | 2 | 0 | fixed |
| summary closers | 2 | 0 | fixed |
| vague attribution | 1 | 0 | fixed |
Diff excerpt. Deletions struck through, insertions on green. This is the hardest text in the set, so a couple of checks are still failing after the rewrite, and the table above says so: honesty over polish.
العربية: منشور تقني، من 74 إلى 98
| الفحص | قبل | بعد | |
|---|---|---|---|
| علامات لاتينية / تطويل | 1 | 0 | أُصلح |
| العبارات الاحتفالية الجاهزة | 4 | 0 | أُصلح |
مقتطف من الفرق. لاحظ الربط بـ«ثم» و«و» بدل الترقيم الآلي:
The whole eval suite, since one example proves nothing
The repo carries a committed evaluation set: 20 texts across genres and lengths in both languages, run end to end against the engine this site actually serves. Latest committed run: mean gain +21.1 points, 14 of 20 texts improved, 0 run errors. The worst English fixture went 16 to 77; the Arabic fixtures finished between 96 and 100. Drafts that failed the faithfulness guard mid-run were discarded before they could win, which is its whole job.
6 texts did not beat their originals, and the tool said exactly that instead of shipping them quietly. That honesty is a feature, not a bug to file.