What a test answers
An A/B test answers one question: does this change get more clients to reply? Example: a custom intro video recorded for each client, against your usual pre-recorded one.
GrrrGrowth keeps it simple. You have at most 2 templates. Turn the test on in Settings, Proposals tab, and new drafts take turns: A, B, A, B.
Change one thing only
Keep both templates the same except for the one thing you are testing. If you change the opening line and the video together, a better result cannot tell you which one did it.
Name each template by its difference, like Custom intro and Generic intro, so the results read plainly.
Templates are locked while the test runs. New text halfway through would mix two versions in one column.
- Same text in both templates
- One difference, written in the name
- Common client questions already answered
- Portfolio and profile done before you start; don't change them mid-test
Let the app take turns
Don't pick which side a job gets. You would give the special version to the jobs you like most, and those reply more anyway. Taking turns keeps both sides fair.
If one side needs extra work, like recording a custom video, do it every time that side comes up. A skipped video makes that side look worse than it is.
How many sends before it means anything
Judge on replies, not wins. Wins are 3 to 5 percent of sends, so telling two sides apart on wins takes around 1,500 sends each. Replies are more common and show sooner.
With a reply rate near 10 percent, here is what each amount of sends per side can show: 20 is noise. 60 shows only a huge gap, like 10 versus 30 percent. 100 shows a big gap, like 10 versus 25. 200 shows 10 versus 20. A small gain, like 10 versus 15, needs about 680 per side.
The results card on Metrics says Too early, Early sign, Strong sign or Solid from the smaller side's sends.
How long that takes
At 4 sends a day, Monday to Friday, each side gets about 10 a week. A big difference shows in 2 to 3 months; a small one never will at that pace.
More sends get you there faster, but only on jobs you would send to anyway. Sending to weak jobs lowers the reply rate on both sides and costs Connects.
What good numbers look like
Upwork publishes no benchmarks, and the public estimates disagree by up to two times. Our rough read of them: replies 5 to 15 percent of sends is typical, under 5 means something is off, 20 or more is strong. Interviews about 10 to 20 percent. Wins 3 to 10 percent, lower for new accounts.
These are averages. Your market can sit far from them: web development, design and writing run at different rates, and so do budget size, how many proposals a job already has, and how established your account is. A new account in a crowded niche can sit well under these ranges and still be doing fine.
So compare against your own numbers over time, not someone else's.
Read where the funnel leaks
Each stage points at a different fix. Look at the first stage that falls short, and fix that before anything later.
- Few proposals viewed: timing, your profile thumbnail, or the first line clients see in the list
- Viewed but few replies: the proposal itself, so this is where an A/B test helps
- Replies but few interviews: how you answer, and how fast
- Interviews but few wins: portfolio, profile proof, or your rate
