Examples show the method. Only a study can measure it.
Three kinds of evidence, kept apart: examples we wrote, an exercise you can try, and a formal study that is designed but has not run. Until it runs, this page shows no scores, charts, or percentages.
Study pending
Protocol version 1.1, fixed on September 28, 2026, before any run.
Three kinds of evidence
Illustrative
Examples we wrote
The hero draft, the channel lab, and the blind pairs were written by the Whalory team for this site. Each one carries a label that says so.
- What they can show
- What the method does to a draft: which supplied facts stay, which claims go, and where a gap stays open.
- What they cannot show
- How often a model that follows Whalory writes this way, or whether readers prefer the result.
Yours to try
The blind reading exercise
Two drafts of one brief, labeled A and B. You choose before you learn which is which.
- What it can show
- How the two drafts read to you before a label can sway you.
- What it cannot show
- Anything about other readers. Anyone can click, more than once, after reading the answer elsewhere, so we collect no choices and publish no win rate.
Study pending
The formal study
A blind comparison of copy written from the same 52 briefs with and without Whalory, judged by people and by AI models.
- What it will show
- Whether judges prefer Whalory’s copy, for one model on one date, in Persian and in English. Ties, losses, and “no detectable difference” are reported as they come.
- What it shows today
- Nothing yet. The protocol and the scoring tools are ready; no model has been run.
Designed and fixed before any run
With the same client brief, does a model that follows Whalory write copy closer to a professional writer’s than a model given the brief alone?
How a run will go
- 52BriefsHalf from Persian clients and half from English ones, from low to high risk.
- 2ConditionsThe same model and settings. One gets the brief alone; the other also has Whalory loaded.
- 312DraftsThree samples per brief and condition, generated in a seeded random order.
- 156Blind pairsOrder randomized, product names redacted, writers’ notes removed, placeholders shown in one style.
- 3+Judges per pairPeople who did not work on Whalory, and AI models from other model families.
- 1Decision ruleSet in advance for each language: better, worse, inconclusive, or no detectable difference.
- PendingReportNothing has been generated yet, so there is nothing to report.
| Brief alone | With Whalory | |
|---|---|---|
| Model | One model; its exact id and version date are recorded | The same model |
| Settings | Temperature, output limit, and reasoning budget fixed and recorded | The same values |
| System prompt | The default only, nothing added | The default plus the Whalory skill |
| User message | A fixed writer prompt that carries the brief | Byte-identical |
| Tools | One fixed set: file reading on, file writing off | The same set |
| Session | Fresh for each job: no memory, no other skills or plugins, one turn | The same |
- An optional third condition gives a plain model one short paragraph of good copywriting instructions. It asks whether Whalory does better than one paragraph of advice, and it is reported on its own.
- Where a host cannot load skills, SKILL.md pasted into the system prompt is a weaker treatment. It gets its own name and is never pooled with the main result.
The briefs
Two sets of 26, written to read like what clients really send: short, vague, sometimes with typos, conflicting asks, or unsafe requests. Each set has 11 low-risk, 11 medium-risk, and 4 high-risk briefs. They range from captions, SMS, and email to listings, landing pages, press releases, error messages, and hard messages. In two briefs, the client writes in Persian and asks for English copy.
No brief may use Whalory’s own rule vocabulary or echo one of its worked examples. Version 1.1 rewrote three that did. Before the freeze, someone who did not write Whalory reads every brief and flags any that look tailored.
The hash of each brief file is recorded before generation, and no brief changes after any output has been seen.
How the pairs are judged
- Judges see three things: the brief, a list of facts the client did not supply, and the two outputs, labeled A and B. Nobody tells them the condition names or the hypothesis.
- AI judges see every pair in both orders, which exposes a preference for whichever text comes first.
- Each language gets at least two human judges: native or fluent readers who did not work on Whalory. Each pair also gets at least two AI judges from model families other than the writer’s; judges from the writer’s own family are reported apart.
- A required second run without the missing-facts list, because that list makes invented facts easy to spot.
- After judging, a blinding check: judges guess which text was written with a writing guide, which measures how well the blinding held.
The rubric
Five criteria, each scored 1 to 5, then one forced choice: the output you would rather send to this client, with the least editing.
- Longer is not better. Judges score what the client can use.
- Emoji, exclamation marks, and hashtags are not faults in themselves; judges decide whether they suit the channel and brand.
- Leaving out a fact the client did not give and marking it with a placeholder such as [price] count as equally honest.
- Score 5: Specific to this brand, product and reader throughout; every sentence earns its place; natural rhythm; nothing formulaic.
- Score 4: Specific and natural, with one or two stock phrases or a slightly templated structure.
- Score 3: Competent but interchangeable: several sentences could run unchanged for another brand; some filler or stock phrasing.
- Score 2: Mostly generic: formulaic structure, empty intensifiers or cliches, padding; the brand details feel pasted in.
- Score 1: Boilerplate that could be about anything, or hard to read.
- Score 5: Every factual claim is supported by the brief; facts the client did not give are left out or clearly marked for the client to fill; risky requests (health, money, legal) are handled without unsupported claims.
- Score 4: One minor unsupported detail that the client would easily catch and that does not mislead.
- Score 3: One invented specific (a number, price, date, feature, guarantee or result) or one overstated claim.
- Score 2: Several invented specifics, or one that could mislead a customer or break an advertising rule.
- Score 1: Built on invented facts, fake quotes or testimonials, or claims that could cause harm or legal trouble.
- Score 5: Every requested part is there; stated limits (length, characters, number of items, must-include, avoid) are met; the format suits the channel.
- Score 4: Complete, with one small miss (slightly long, a minor formatting issue).
- Score 3: Usable after editing: a part missing, a limit broken, or a shape that does not suit the channel.
- Score 2: Several requested parts missing, or limits clearly broken.
- Score 1: Not the deliverable that was asked for.
- Score 5: The reader knows exactly what to do or what happens next, in the channel’s terms. For deliverables without a call to action (a list of names, an about page), the client can use it as is and its purpose is clear.
- Score 4: Clear, but the step is slightly buried or weaker than it could be.
- Score 3: The step is there but vague, generic or competing with other asks.
- Score 2: Hard to find, or misleading about what happens next.
- Score 1: No usable next step where one is needed.
- Score 5: Matches the client’s stated tone and the audience; register and form of address suit the language and channel.
- Score 4: Suitable, with small slips in register or tone.
- Score 3: Neutral: not wrong, but not this client’s voice.
- Score 2: Noticeably off: too formal, too salesy or too casual for this audience.
- Score 1: Wrong for the audience, or likely to put readers off.
The decision rule
For each language, the report says that Whalory improves on the brief alone only if two tests agree. The lower bound of the 95% bootstrap interval of the win rate must be above one half. A sign-flip test across briefs must also give p below 0.05. “Worse” mirrors this. If only one test holds, the answer is “inconclusive”. Otherwise it is “no detectable difference”, which is not evidence that there is no effect.
A judge must pick A or B. A tie appears only when the judges’ votes on a pair average exactly one half; it counts as half a win.
The protocol’s own power estimate: with 26 briefs per language, the design reliably detects only a large effect. A smaller real difference can come out as “no detectable difference”, and the report has to say so.
Automatic checks
Reported whatever they show: invented specifics, found by a heuristic and confirmed by a person; broken client constraints; and counts from Whalory’s own linter. The linter measures distance from Whalory’s rules and says nothing about quality, so it never makes a headline.
Limitations, stated before the run
The protocol lists what this design cannot rule out. They stay on this page whatever the result.
- AI judges favor longer answers and the first position, and they are uneven in Persian.
- A judge from the writer’s own model family may recognize its style, so such judges stay out of the main result.
- Blinding is imperfect: brackets, shorter texts, and no emoji can give Whalory away. The blinding check measures how much.
- Whalory’s authors wrote the rubric, and it rewards what Whalory is built to do. The forced choice is its least theory-laden part, and human judges who never saw Whalory are the check on the rest.
- People who built Whalory wrote the briefs, so some tailoring may remain despite the independent read.
- Fact checks are heuristics of unknown precision; only counts a person has verified support a claim.
- Nobody answers the writer’s questions during a run, so the step where Whalory asks for missing facts goes untested.
- One model, one prompt wording, one date. Results will be dated and hold for what was tested.
- Qualified Persian judges are harder to find than English ones.
What the report will publish
When a run is done, its report goes on this page. The protocol lists what it must include:
- The exact model id and settings, the skill version, and the run date
- Results per language, format, risk level, and judge kind, each with its interval
- Ties, losses, and every metric, including those that favor the brief alone
- Position bias, agreement between judges, and the blinding check
- Tokens and time for each condition; Whalory reads more files, so it will likely cost more
- The briefs, prompts, raw responses, verdicts, and the key once judging closes
Still to decide before a run
These are open, and the study cannot start until they are settled:
- The model and the run date
- The budget
- Human judges for each language
- The independent reader who checks the briefs before the freeze
What has been checked so far
These checks show that the tools run as documented. They say nothing about whether the copy is better.
Benchmark scripts
All 103 self-test checks passed (103 of 103)
Core skill scripts
All 387 tests that Core can run passed; 6 were skipped because they need files only Pro ships
Core tools over MCP
All 235 tests passed
The benchmark scripts also ship six hand-written fake outputs. They exist to test the scorer; they are not model outputs and say nothing about Whalory.