Evidence

Examples show the method. Only a study can measure it.

Three kinds of evidence, kept apart: examples we wrote, an exercise you can try, and a formal study that is designed but has not run. Until it runs, this page shows no scores, charts, or percentages.

Study pending

Protocol version 1.1, fixed on September 28, 2026, before any run.

Three kinds of evidence

  1. Illustrative

    Examples we wrote

    The hero draft, the channel lab, and the blind pairs were written by the Whalory team for this site. Each one carries a label that says so.

    What they can show
    What the method does to a draft: which supplied facts stay, which claims go, and where a gap stays open.
    What they cannot show
    How often a model that follows Whalory writes this way, or whether readers prefer the result.
  2. Yours to try

    The blind reading exercise

    Two drafts of one brief, labeled A and B. You choose before you learn which is which.

    What it can show
    How the two drafts read to you before a label can sway you.
    What it cannot show
    Anything about other readers. Anyone can click, more than once, after reading the answer elsewhere, so we collect no choices and publish no win rate.
  3. Study pending

    The formal study

    A blind comparison of copy written from the same 52 briefs with and without Whalory, judged by people and by AI models.

    What it will show
    Whether judges prefer Whalory’s copy, for one model on one date, in Persian and in English. Ties, losses, and “no detectable difference” are reported as they come.
    What it shows today
    Nothing yet. The protocol and the scoring tools are ready; no model has been run.

The formal study

Designed and fixed before any run

The question

With the same client brief, does a model that follows Whalory write copy closer to a professional writer’s than a model given the brief alone?

How a run will go

  1. 52BriefsHalf from Persian clients and half from English ones, from low to high risk.
  2. 2ConditionsThe same model and settings. One gets the brief alone; the other also has Whalory loaded.
  3. 312DraftsThree samples per brief and condition, generated in a seeded random order.
  4. 156Blind pairsOrder randomized, product names redacted, writers’ notes removed, placeholders shown in one style.
  5. 3+Judges per pairPeople who did not work on Whalory, and AI models from other model families.
  6. 1Decision ruleSet in advance for each language: better, worse, inconclusive, or no detectable difference.
  7. PendingReportNothing has been generated yet, so there is nothing to report.
What changes between the two conditionsEverything is held equal except the skill.
Brief aloneWith Whalory
ModelOne model; its exact id and version date are recordedThe same model
SettingsTemperature, output limit, and reasoning budget fixed and recordedThe same values
System promptThe default only, nothing addedThe default plus the Whalory skill
User messageA fixed writer prompt that carries the briefByte-identical
ToolsOne fixed set: file reading on, file writing offThe same set
SessionFresh for each job: no memory, no other skills or plugins, one turnThe same
  • An optional third condition gives a plain model one short paragraph of good copywriting instructions. It asks whether Whalory does better than one paragraph of advice, and it is reported on its own.
  • Where a host cannot load skills, SKILL.md pasted into the system prompt is a weaker treatment. It gets its own name and is never pooled with the main result.

The briefs

Two sets of 26, written to read like what clients really send: short, vague, sometimes with typos, conflicting asks, or unsafe requests. Each set has 11 low-risk, 11 medium-risk, and 4 high-risk briefs. They range from captions, SMS, and email to listings, landing pages, press releases, error messages, and hard messages. In two briefs, the client writes in Persian and asks for English copy.

No brief may use Whalory’s own rule vocabulary or echo one of its worked examples. Version 1.1 rewrote three that did. Before the freeze, someone who did not write Whalory reads every brief and flags any that look tailored.

The hash of each brief file is recorded before generation, and no brief changes after any output has been seen.

How the pairs are judged

  • Judges see three things: the brief, a list of facts the client did not supply, and the two outputs, labeled A and B. Nobody tells them the condition names or the hypothesis.
  • AI judges see every pair in both orders, which exposes a preference for whichever text comes first.
  • Each language gets at least two human judges: native or fluent readers who did not work on Whalory. Each pair also gets at least two AI judges from model families other than the writer’s; judges from the writer’s own family are reported apart.
  • A required second run without the missing-facts list, because that list makes invented facts easy to spot.
  • After judging, a blinding check: judges guess which text was written with a writing guide, which measures how well the blinding held.

The rubric

Five criteria, each scored 1 to 5, then one forced choice: the output you would rather send to this client, with the least editing.

  • Longer is not better. Judges score what the client can use.
  • Emoji, exclamation marks, and hashtags are not faults in themselves; judges decide whether they suit the channel and brand.
  • Leaving out a fact the client did not give and marking it with a placeholder such as [price] count as equally honest.

The decision rule

For each language, the report says that Whalory improves on the brief alone only if two tests agree. The lower bound of the 95% bootstrap interval of the win rate must be above one half. A sign-flip test across briefs must also give p below 0.05. “Worse” mirrors this. If only one test holds, the answer is “inconclusive”. Otherwise it is “no detectable difference”, which is not evidence that there is no effect.

A judge must pick A or B. A tie appears only when the judges’ votes on a pair average exactly one half; it counts as half a win.

The protocol’s own power estimate: with 26 briefs per language, the design reliably detects only a large effect. A smaller real difference can come out as “no detectable difference”, and the report has to say so.

Automatic checks

Reported whatever they show: invented specifics, found by a heuristic and confirmed by a person; broken client constraints; and counts from Whalory’s own linter. The linter measures distance from Whalory’s rules and says nothing about quality, so it never makes a headline.

Limitations, stated before the run

The protocol lists what this design cannot rule out. They stay on this page whatever the result.

  1. AI judges favor longer answers and the first position, and they are uneven in Persian.
  2. A judge from the writer’s own model family may recognize its style, so such judges stay out of the main result.
  3. Blinding is imperfect: brackets, shorter texts, and no emoji can give Whalory away. The blinding check measures how much.
  4. Whalory’s authors wrote the rubric, and it rewards what Whalory is built to do. The forced choice is its least theory-laden part, and human judges who never saw Whalory are the check on the rest.
  5. People who built Whalory wrote the briefs, so some tailoring may remain despite the independent read.
  6. Fact checks are heuristics of unknown precision; only counts a person has verified support a claim.
  7. Nobody answers the writer’s questions during a run, so the step where Whalory asks for missing facts goes untested.
  8. One model, one prompt wording, one date. Results will be dated and hold for what was tested.
  9. Qualified Persian judges are harder to find than English ones.

What the report will publish

When a run is done, its report goes on this page. The protocol lists what it must include:

  • The exact model id and settings, the skill version, and the run date
  • Results per language, format, risk level, and judge kind, each with its interval
  • Ties, losses, and every metric, including those that favor the brief alone
  • Position bias, agreement between judges, and the blinding check
  • Tokens and time for each condition; Whalory reads more files, so it will likely cost more
  • The briefs, prompts, raw responses, verdicts, and the key once judging closes

Still to decide before a run

These are open, and the study cannot start until they are settled:

  • The model and the run date
  • The budget
  • Human judges for each language
  • The independent reader who checks the briefs before the freeze

What has been checked so far

These checks show that the tools run as documented. They say nothing about whether the copy is better.

  • Benchmark scripts

    All 103 self-test checks passed (103 of 103)

    python bench/selftest_bench.pySeptember 28, 2026

  • Core skill scripts

    All 387 tests that Core can run passed; 6 were skipped because they need files only Pro ships

    python scripts/selftest.pySeptember 28, 2026

  • Core tools over MCP

    All 235 tests passed

    python scripts/selftest_mcp.pySeptember 28, 2026

The benchmark scripts also ship six hand-written fake outputs. They exist to test the scorer; they are not model outputs and say nothing about Whalory.

Read two drafts before the labels