One test tells you whether your own inbound replies were written by a person. It takes about ten minutes on data you already hold, and it is not the test your team is being scored on.
Every system that grades an inbound reply grades its manner. Did it use the buyer's first name. Did it come from a named person rather than a noreply address. Did it ask a question. Did it arrive quickly. Reply scoring inside a sales engagement platform works this way, and so does the manager who reads a handful of emails and decides the team sounds attentive.
Manner is the first thing software learned to copy, and the cheapest. None of it requires knowing anything about the buyer that the buyer did not just type into a form.
So we went looking for what actually separates a reply a person wrote from a reply a system sent, on emails where we already knew the answer. Across 2023 we asked for a demo on B2B websites, posing as a mid-sized US software company, and captured everything that came back. Then we hand classified a sample of the replies one at a time, reading each in full.
The separator was not a way of writing. It was an act. In almost every reply a person had written, the sender had left the form, done something, and then said what they found.
This study has no outcome data. It records what each company sent and when, and nothing about what happened next. Not one of these emails is joined to a result.
So nothing here can tell you what works. It can tell you what the few who put a person in front of the buyer did differently from the many who did not, and that comes down to one thing you can check yourself. The fieldwork ran in 2023, so every rate here is a baseline rather than a current reading.
The findings
| Feature | Human replies, base 17 | Acknowledgments, base 162 |
|---|---|---|
| An act performed on this submission, with the fact it turned up named | 14 | 5 |
| A phone call or voicemail reported | 4 | 1 |
| A fact about the buyer that had to be looked up or worked out | 7 | 3 |
| A fit verdict rendered, usually a refusal | 4 | 3 |
| A question asked | 12, 71% | 115, 71% |
| "Thank you for your interest" or a variant | 3, 18% | 27, 17% |
| A named personal sender address | 17 | 152 |
| A booking link | 7, 41% | 61, 38% |
| The buyer's company name used in the body | 4, 24% | 28, 17% |
| Median reply lag | 765 minutes | 281 minutes |
Source: 190 emails hand classified from the 2023 demo response study, returning 17 human replies, 162 acknowledgments and 11 marketing blasts. The blasts are left out of the two columns because they separate on length and link count alone. Across the wider study, 93.4% of the 1,685 companies whose form we could submit never put a human in front of the buyer.
1. Somebody had gone and done something first
The sender reported an act they had personally performed on this specific submission, and named what it turned up. Present in 14 of the 17 human replies and 5 of the 162 acknowledgments.
Nothing else in the feature set came close. It is almost always there in the human class and almost never anywhere else.
The buyer could tell without being told. The reply came back carrying something he had never typed into the form. A call he had missed. A line off his own website. A fact about his company that no field had asked for.
Four kinds of act account for all fourteen.
Someone picked up the phone, 4 of the 17. A property management platform wrote, "I tried calling again just now and left a message. I reached out yesterday but wasn't able to connect." An incident management vendor opened by reporting the voicemail he had already left. He explained why his company runs a short call before a demo, then offered eleven named times across three days.
The voicemail is what classified it. A list of times is something scheduling software produces on its own. The shortest human reply in the sample, 233 characters with no signature block and no links, was a rep reporting a call attempt.
Someone went and looked, 7 of the 17. The chief executive of a small analytics company said he would bring his VP of Technology to the call, "He's a former Marine like you (I saw your LI profile), and I'm a Navy vet." A customer engagement vendor wrote that it had looked over the buyer's website. It correctly named the category the company sells into.
A conversational messaging vendor read the buyer's email domain, decided the form's stated industry was wrong, and said so: "What type of business do you have? From the email provided, it looks like you're a software?"
Someone disclosed the state of their own side, 2 of the 17. An identity vendor told the buyer the product he had asked about was being sunset. It named the parent company that had bought it, and offered a named colleague as a faster route. A chief operating officer said demos are run out of the Nashville office and asked whether the buyer would be at a named trade show.
Someone checked and could not help, 1 of the 17. An accounting platform told the buyer the demo he wanted did not exist in his language, which no other requester would have been told.
On the other side of the line, 5 of the 162 acknowledgments report an act of any kind, and exactly one of those reports a phone call the sender had placed.
2. Somebody was willing to tell him no
A fit verdict, 4 of 17 human replies against 3 of 162 acknowledgments. Small numbers on both sides, and the ratio is 24% against 2%.
A refusal is not a welcome email, and it is still an answer. The buyer who got one knew where he stood that day. He knew somebody had held his business up against the seller's and made a call on it. The buyer who got the named rep first touch knew neither.
In three of those four the verdict went against the buyer, and in two of those three it closed the door. A brand analytics vendor declined and published the bar it was applying: consumer brands above a stated revenue line. The buyer was not one.
A lab software vendor opened with "I'm curious about your interest as your company doesn't fit the mold based on who we typically work with (pharma, biotech, etc.)" Then it offered three call lengths anyway. An accounting platform in a non-English market refused outright, having no demo in the buyer's language. It sent a subtitled tutorial and a phone number instead.
The largest single group inside the 162 acknowledgments is the named rep first touch, 98 of them. Not one told the buyer no. A refusal costs the sender something and requires a judgment about this buyer against this business, which is why templates do not contain them.
3. The same email everybody else got
Five patterns account for all 162 acknowledgments. A named rep first touch, 98. A qualifying questionnaire, two or more questions the buyer must answer before anything happens, 45. An auto receipt promising that someone will be in touch, 11. A follow-up bump on a previous template, 5. The buyer's own form submission bounced back to him, 3.
Set that against what the buyer did. He found the company, decided to talk, and filled in the form. The most common thing that came back was a person's name attached to a paragraph. The same paragraph would have gone, word for word, to anyone else who filled in the same form.
4. Six signals that separated nothing
This is where most reply quality scoring lives, and every one of these is measured on the same 17 human replies and 162 acknowledgments.
Asking a question. 12 of 17 human replies contained a question mark, and 115 of 162 acknowledgments did. The same 71% in both classes. If anything it points the wrong way, because the 45 qualifying questionnaires ask more questions than the average human reply does. Interrogation scales and conversation does not.
"Thank you for your interest." 3 of 17 human replies, 27 of 162 acknowledgments. 18% against 17%. The phrase every guide tells you to strip out of your templates appears at the same rate in the emails people actually wrote.
A named personal sender address. 17 of 17 human replies, 152 of 162 acknowledgments. Inside the groups where the question is actually decided, a real name in the from field is a constant, not a variable. A role address was never human in this sample, so its presence is a useful exclusion, but its absence tells you nothing at all.
Speed. It runs backwards. Median reply lag was 765 minutes for the emails a person wrote and 281 minutes for automated acknowledgments, on the same bases of 17 and 162.
Replies landing inside an hour: 4 of 17 human against 54 of 162 acknowledgments. The three fastest replies in the entire hand classified set arrived at 0 seconds, 3 seconds and 5 seconds. All three were machines.
Every hour you take off your response time is an hour taken off the part of the process that a person was never in. The same argument, worked through on the whole corpus, is in Speed to Lead is timing a robot.
A booking link. 7 of 17 human replies, 61 of 162 acknowledgments. 41% against 38%. Scheduling software emits calendar links, and so do people.
The buyer's company name in the body. Present in 4 of 17 human replies and 28 of 162 acknowledgments, 24% against 17%. Length behaves the same way: median 692 characters for the human replies against 699 for the acknowledgments.
Those six features are close to an exact description of what an automated classifier weights. So we built one, scored it against the hand verdicts, and threw it away. On all 190 emails it agreed with the hand answer 70 times. It called 116 emails human where the correct count was 17, a precision of 9%.
Train a machine on what the industry believes a personal reply looks like and it calls the well-written template a person. The industry's picture of a personal reply is a picture of a good template.
5. Merge fields prove nothing
Every submission in this study came from one buyer identity, and no form we filled in offered a free text message field. The identity did carry a LinkedIn profile and the company did have a website. That is why a reply reporting that someone read either one counts as an act: it could not have come out of a merge field. So the first name in the reply, the company name, the domain, the industry, the headcount: all of it came back out of the same form it went into. An email using those tokens is reproducing data the buyer supplied a few minutes earlier.
That is what makes the near misses instructive. Five acknowledgments in the 162 do report an act. The strongest of them opens by saying the sender had just tried to call, in French, before continuing into a fixed pack of company press links. Another refuses a demo because the integration the buyer needs has not shipped yet. That is real information. It is also derivable by rule from one form field, so every buyer on that platform gets the same email.
A third sounds closest of all and gets nowhere. It opens "Looking at your requirements", then names nothing it found and asks four generic questions.
The line falls in the same place every time. Specificity a rule could generate from a form field stayed an acknowledgment. Specificity that required leaving the form was a person.
6. The person usually arrived second
10 of the 17 human replies were not the first thing the company sent. They were the second email, arriving after something else.
We went back and read the first email in each of those cases. In 6 of the 10 it had been an auto receipt or a bulk send. The person appeared only on the follow-up, usually from a different address. Across the whole study, 984 of the 1,685 companies sent exactly one email and never a second, 58.4%, inside a capture that stopped at two.
The buyer met the machine first. Whatever impression the company made in that opening email, it made with nobody there.
Any measurement that looks only at the first response is looking at the position where the machine sits.
7. What to check on your own last hundred replies
Ten minutes, on data you already hold.
- Pull your last hundred first replies to inbound demo or contact requests, out of sent items rather than out of a CRM activity summary.
- Delete every merge field before you judge anything. First name, company name, domain, industry, headcount. All of it came from the form.
- Now ask one question of what remains. Does this email state something the sender personally did about this submission before writing it, and name what it turned up? A call placed. A voicemail left. A page read. A profile checked. An odd field answer queried.
- Count how many survive. That count, over a hundred, is your human reply rate.
- Repeat it on the second email in each sequence, not just the first. In our data that is where the person usually was.
It will not tell you whether those replies produced meetings. We have no outcome data, and neither does a count of your own sent items. It will not tell you whether the team is trying, because the reps writing the templated first touches are usually working exactly as instructed.
It tells you how many of your buyers heard from a person at all. Most B2B companies would answer that with a number close to zero. Of the 1,685 companies we approached, 93.4% never put a human in front of the buyer. An estimated 6.6% did.
8. Method
2,528 B2B companies, all with more than 10,000 monthly website visitors, were sent one demo or contact sales request each during 2023. 798 forms could not be submitted and 45 errored, leaving a usable base of 1,685 companies and 2,386 captured emails, of which 2,102 have a readable body.
Whether a person wrote an email was decided by hand, one email at a time. Each verdict was made from the full recovered body rather than from a stored preview. 190 emails were classified across two independent batches drawn group by group. The bulk infrastructure group was sampled at 30 of 1,060 population emails. The personal sender with no ask group at 35 of 186. And the deciding group, personal sender who asks for something, at 125 of 856.
That group was drawn twice and returned 12.3% human on the first draw and 11.7% on the second. The published rates are scaled back up from each group to the full 1,685, in proportion to how many emails each group actually held. The raw 8.9% human share of the 190 is not a rate for the market: it reflects how the sample was built, which deliberately over-read the group where the humans were.
The population rates carry these bounds. 1,573 of the 1,685 companies never put a human in front of the buyer, which is the 93.4%, and its 95% interval spans 90.3% to 96.1%. The 6.6% who did put a human in front of the buyer works out at about 112 companies, on a 95% interval from 3.9% to 9.7%.
The 11 marketing blasts held out of the comparison table separate on length and link count alone. Median 2,844 characters and 5 embedded links, against 692 characters and a median of no links at all in the human replies.
Two limits govern how far these figures stretch. The human class is 17 emails, so every proportion quoted for it is directional and a single reclassification moves it by about six points. And collection captured at most two emails per company. Human contact more often arrived at the second email than the first. So the true share of companies that never put a human in front of the buyer may be about a point lower than 93.4%, and no more than 2.4 points lower.
Full method is on the demo response study page, and the group weights and classification rules are on the classification rulebook page. Every rate quoted here, with its base, sits in the full figure set. The demo response study is reported in full in The Buyer Wait Time Report 2026.
Two earlier audits ran this design and counted arrivals, the 2011 Harvard Business Review one and Drift's in 2017. Neither opened the emails to ask who wrote them, which is the variable this article is about. The lineage is on the demo response study page.
9. Questions
Our reps write the first line themselves. Does that count as a human reply?
Apply the test to it. If the line is built from something on the form, it is an acknowledgment however it was typed. The same line comes out for the next buyer whose form says the same thing.
If it reports that the rep called, or read something, or found something the form did not supply, it counts. Only 5 of our 162 acknowledgments cleared that bar, and in four of those five the specific detail was derivable by rule from a form field. The fifth reported a phone call the sender had placed.
Are you saying response speed does not matter?
I am saying it does not tell you whether a person is there, which is what it is currently used for. Median reply lag ran to 765 minutes for the 17 emails a person wrote, against 281 minutes for the 162 acknowledgments. The three fastest replies in the whole set were machines, answering in 0, 3 and 5 seconds.
Whether a slower human reply beats a faster machine one on any commercial measure is a question this study cannot answer, because it has no outcome data.
Seventeen emails is a very small base.
It is, and every percentage quoted for the human class should be read as directional. Two things hold it up. The design put the sample where the humans were, in the group where the question is actually decided. That group was drawn twice and landed 0.6 points apart.
And the separating feature does not need precision to survive. 14 of 17 against 5 of 162 is not a gap that a reclassification or two closes.
Is it worth answering this way, given what it costs?
I cannot tell you it pays, and I would not trust anyone who claims to know it from this data.
What I can tell you is the rate. Better than nine in ten of the 1,685 companies never got a person in front of a buyer who had asked to see the product. And of the 162 acknowledgments we read, 98 were a named salesperson introducing themselves without one fact from outside the form.
The fieldwork is from 2023. Is it still current?
Treat the rates as a baseline. Everything since has made manner easier to fake, through better generated prose and cheaper personalization tokens, while leaving the act itself exactly as expensive as it was. I have not measured that, and nobody has.
Terry Wilson is the founder of GTM Clarity and CEO of ChatMetrics, which has delivered over $5 billion in qualified pipeline and 300,000+ leads for B2B clients across SaaS, services, and industrial sectors. Before founding ChatMetrics, Terry was National Sales & Marketing Manager for a $1B enterprise, leading more than 350 people across Australia. He built GTM Clarity's AI on a corpus of 3M+ real B2B sales conversations that delivered $5B+ pipeline across 200+ companies.
Keep reading
The Buyer Wait Time Report 2026
Think about what has to happen before a person fills in your demo form.
Read →Statistics62 B2B buyer response statistics, from five years of original research
Base for this section: 4,265 B2B technology websites, each above 10,000 monthly visitors, classified in 2021, again in 2025 and again in 2026 at the same addresses. The 2021 and 2025 figures are observed states with no coder judgment about intent, no estimation and no sampling. The 2026 wave sampled the three big groups and tried every site in the three small ones, so its figures carry confidence intervals and the first two waves do not. It is also a different instrument: 2021 and 2025 classified a site by opening the widget on one visit, while 2026 held a conversation and waited three minutes for a person.
Read →ResearchSpeed to Lead is timing a robot
Every B2B sales team runs on a stopwatch. The clock starts when a buyer raises a hand and stops when the company responds.
Read →