# We Tested 7 iMessage Copy Rules on 45,217 Real Messages. Six Failed.

> What actually moves reply rate — and what turns out to be one customer's house style

*By BG · Published 2026-09-11T00:00:00.000Z*

**Canonical URL:** https://tuco.ai/blog/imessage-copy-rules-tested-45217-messages
**Tags:** iMessage benchmarks, text message templates, reply rate, original research

## Summary

We ran 7 common text-message copy rules against 21,305 first-touch iMessages from 48 accounts. Only one survived a within-account test, and even that one is a coin flip. Who you message matters 35x more than how you word it.

---


**We analysed 45,217 outbound iMessages and could not find a copy rule that holds.** Seven common pieces of texting advice, tested against 21,305 first-touch messages from 48 accounts. Six failed. The seventh is a coin flip. Meanwhile the same data shows a **65.5 percentage point** spread between accounts — meaning who you message matters roughly 35 times more than how you word it.

This is our own production data, and the method is below so you can argue with it.

## The short answer

| Copy rule | Verdict |
|---|---|
| Shorter messages get more replies | **Failed** — 61% of the short messages came from one account |
| End with a question | **Failed** — -1.7pp inside accounts, helped in 5 of 10 |
| Use an exclamation mark | **Failed** — 68% of one account |
| Open with "Hi/Hey {name}" | **Failed** — the no-greeting arm was 79% one account |
| Avoid numbers and times | **Failed** — 54% one account |
| Keep it under 20 words | **Failed** — 67% one account |
| Single-line beats multi-line | **Survived pooling, then failed** — better in 6 accounts, worse in 5 |

## Why most "copy rules" are one customer's house style

Here is the trap, and it is the reason this article exists rather than a list of templates.

Pool everyone's messages together and patterns appear immediately. Single-line messages replied at **17.8%** against **8.9%** for multi-line. That is a doubling, on 21,305 messages. It looks like a finding.

It is not. It is [Simpson's paradox](https://en.wikipedia.org/wiki/Simpson%27s_paradox). Accounts that happen to write single-line messages are accounts that happen to get replies — cash-pay clinics texting their own patients — and accounts writing long multi-block messages skew toward cold B2B prospecting. The copy is a marker for the audience, not a cause of the reply.

The test that settles it: **measure inside one account**, where the audience is held constant.

Twelve accounts sent at least 30 of both single-line and multi-line first-touch messages. Single-line replied better in **6**, worse in **5**, tied in **1**. Mean difference **+0.9 percentage points**, range **-30.2 to +21.7**.

That is a coin flip with a wide error bar. We are not going to sell you a template library on the back of it.

## The seven rules, tested

Every arm below had to clear the same gate before we would even report it: at least 300 messages, at least 15 accounts, and no single account contributing more than 40%. Most did not clear it.

### 1. "Keep it short" — failed

Over a 90-day window, messages under 60 characters replied at **36.4%** against roughly 17% for everything else, spread across 38 accounts with no account above 27%. That looked solid enough that we nearly published it.

Widening to the full corpus doubled the sample and broke it: the under-60 bucket became **61% one account**. The effect was that account, not the length.

### 2. "End with a question" — failed

Pooled: **15.6%** with a question versus **9.3%** without. Inside accounts: mean **-1.7pp**, positive in **5 of 10**. Range -24.9 to +30.5. Nothing there.

### 3. "Use an exclamation mark" — failed the gate

Pooled it looks positive — 15.4% versus 10.5%. But **68% of the exclamation arm is one account**. Untestable.

### 4. "Open with a greeting" — failed the gate

Messages opening "Hi", "Hey" or "Hello" replied at **15.8%** against **4.1%** for no greeting. An enormous gap, and entirely unusable: the no-greeting arm is **79% one account**. That account is doing something else wrong.

### 5. "Avoid numbers and times" — failed the gate

Messages containing a digit replied at **6.7%** versus **16.8%** without. The digit arm is **54% one account** — plausibly one sender pushing appointment times at a cold list.

### 6. "Keep it under 20 words" — failed the gate, and inverted

Under 20 words replied at **6.5%**; over 20 words at **15.4%**. The opposite of the folk wisdom — but the short arm is **67% one account**, so neither direction is trustworthy.

### 7. "Single-line beats multi-line" — survived pooling, failed the paired test

Covered above. The only rule that cleared the distribution gate, and it evaporated when measured inside accounts.

## What actually predicts reply rate

The audience. It is not close.

Across 20 accounts sending at least 100 first-touch messages each:

| | Reply rate |
|---|---|
| Lowest account | **0.3%** |
| Median account | **24.4%** |
| Highest account | **65.8%** |
| **Spread** | **65.5 percentage points** |

Against that, every copy variable we could test honestly moved reply rate by **under 2 percentage points on average** inside an account.

That ratio — roughly 35 to 1 — is the finding. It also matches [our vertical benchmarks](/imessage-benchmarks): cash-pay clinics reply at 54.5%, coaching and webinars at 34.2%, home services at 24.6%, and cold B2B prospecting at 4.1%. No wording closes a gap that size.

## What we could not test

**Emoji.** Only 307 of 21,305 first-touch messages contained one, and **91% came from a single account**. We are reporting this as untestable rather than publishing a number that describes one customer.

**Links.** 89 messages. Far too few.

**Personalisation tokens.** Merge fields are resolved before send, so by the time a message is stored we cannot distinguish a typed first name from a merged one without reading content — which we deliberately did not do.

**Send time.** Two hours showed dramatic differences (9.4% and 9.5% against a 19.6% baseline) and both were 55–69% one account. Bulk sends, not timing effects.

## Method

- **Corpus:** 45,217 outbound iMessages sent from the Tuco AI platform since 15 December 2025.
- **Unit of analysis:** the first outbound iMessage per lead — 21,305 messages across 48 accounts. First touch, because reply attribution on later messages is ambiguous.
- **Outcome:** a lead counts as replied if any inbound message exists for that lead. Reply rate only.
- **Features** were computed inside the database aggregation — character length, line count, word count, punctuation flags, greeting pattern, send hour. **No message content was exported**, and no contact names, phone numbers or email addresses were read at any point.
- **Concentration gate:** every reported arm required n ≥ 300, ≥ 15 accounts, and no account above 40%.
- **Paired test:** for the surviving rule, reply rates were compared within each account that had ≥ 30 messages in both arms.

### Limitations, stated plainly

Reply rate is not a booking rate — nothing in our platform observes whether a reply became an appointment, and any vendor quoting you a conversion rate for this channel is quoting something they cannot see. These are businesses that already chose iMessage and run opt-in lists, so the sample skews positive against outreach generally. Account is a proxy for audience, not a perfect one. And the paired test has 10–12 accounts, which is enough to reject a strong effect but not to detect a small one.

[Download the vertical dataset (CSV, CC BY 4.0)](/imessage-benchmarks/reply-rate-by-vertical.csv) — per-row dated, caveats included in the data.

## So what should you actually do

Stop optimising the words and fix the list. On this data the highest-leverage decisions, in order, are:

1. **Message people who asked to hear from you.** Cold, purchased lists reply at 4.1% and get numbers banned. We [refuse that use](/limitations) as policy.
2. **Pick the audience deliberately.** The gap between the best and worst account is 65 points. No template closes it.
3. **Send fast.** Coaching and webinar audiences reply at a 13-minute median. Speed is a property of the system, not the copy.
4. **Then write like a person.** Once the first three are right, the wording is largely free to be whatever sounds like you — which is what this data actually shows.

If you want the wording anyway, [our templates are free and ungated](/templates). We just are not going to claim they are why the message worked.

**Related:** [Measured reply rates by vertical](/imessage-benchmarks) · [What Tuco AI cannot do](/limitations) · [How we rank every platform in this category](/best/imessage-platforms)


## Frequently Asked Questions

### Do shorter text messages get more replies?

Not reliably. In our data, first-touch iMessages under 60 characters replied at 36.4% over a 90-day window versus about 17% for longer ones — but when we widened to the full corpus, 61% of those short messages came from a single account, and the effect collapsed. Measured inside individual accounts, message length has a mean effect of under 1 percentage point on reply rate and points in different directions for different senders. Length is not a reliable lever.

### Should you end a sales text with a question?

Our data says it makes almost no difference. Pooled across 21,305 first-touch iMessages, messages containing a question replied at 15.6% versus 9.3% without — but that gap is a composition effect. Measured within individual accounts, the mean difference is -1.7 percentage points and questions helped in only 5 of 10 accounts that sent both kinds. If ending with a question worked, it would work inside an account, not just across them.

### Do emojis increase text message reply rates?

We could not test it honestly. Only 307 of 21,305 first-touch messages contained an emoji, and 91% of those came from a single account. Any number we published would describe that one customer's style rather than emoji. We are reporting it as untestable rather than guessing.

### What actually predicts iMessage reply rate?

The audience, by a wide margin. Across 20 accounts sending at least 100 first-touch messages each, reply rates range from 0.3% to 65.8% — a 65.5 percentage point spread. Every copy variable we tested moved reply rate by under 2 percentage points on average when measured inside an account. Cash-pay clinics reply at 54.5% and cold B2B prospecting at 4.1%, and no wording closes that gap.

### How was this measured?

45,217 outbound iMessages sent from the Tuco AI platform since 15 December 2025, narrowed to the first outbound iMessage per lead to give 21,305 messages across 48 accounts. A lead counts as replied if any inbound message exists on that lead. Features were computed inside the database query — character length, line count, word count, punctuation, greeting pattern — so no message content was exported. Reply rate only: bookings happen in the customer's CRM, not ours, so we cannot and do not report conversion.
