offbeat beta
Examples

Real runs, real numbers.

The scores and tables below are read straight from the committed eval run. They update themselves when the engine is re-measured, so the page can never quietly disagree with the numbers it ships.

English: a productivity blog post, score 16 to 77

1677
hard checks 9 to 2 | worst fixture in the set
checkbeforeafter
sentence-length stdev3.31.3still failing
em dashes11still failing
AI-inflation vocabulary120fixed
stock AI phrases20fixed
negative parallelism10fixed
participle tails30fixed
significance inflation20fixed
summary closers20fixed
vague attribution10fixed

Diff excerpt. Deletions struck through, insertions on green. This is the hardest text in the set, so a couple of checks are still failing after the rewrite, and the table above says so: honesty over polish.

In today's rapidly evolving digital landscape, productivity Productivity tools have become pivotal for knowledge workers. It's not just about managing tasks, it's about creating a seamless workflow that empowers matter more than most teams realize. Managing tasks is the easy part. What matters is a workflow that lets teams to thrive. Moreover, the right tool serves as a testament to The right tool reflects a company's culture.

العربية: منشور تقني، من 74 إلى 98

7498
hard checks 2 to 0 | protected content untouched
الفحصقبلبعد
علامات لاتينية / تطويل10أُصلح
العبارات الاحتفالية الجاهزة40أُصلح

مقتطف من الفرق. لاحظ الربط بـ«ثم» و«و» بدل الترقيم الآلي:

أولًا، قمنا بدأنا بترحيل الملفات القديمة. ثانيًا، قمنا بتوحيد القديمة، ثم وحّدنا التنسيق عبر جميع الصفحات. ثالثًا، قمنا بتفعيل الصفحات، وفعّلنا بعدها المعاينة التلقائية لكل مساهمة جديدة.

The whole eval suite, since one example proves nothing

The repo carries a committed evaluation set: 20 texts across genres and lengths in both languages, run end to end against the engine this site actually serves. Latest committed run: mean gain +21.1 points, 14 of 20 texts improved, 0 run errors. The worst English fixture went 16 to 77; the Arabic fixtures finished between 96 and 100. Drafts that failed the faithfulness guard mid-run were discarded before they could win, which is its whole job.

6 texts did not beat their originals, and the tool said exactly that instead of shipping them quietly. That honesty is a feature, not a bug to file.

Run your own text