Full figure set: 62 B2B buyer response statistics.
The published human reply rate comes from 190 emails classified by hand, one at a time, by Terry Wilson, across two independent stratified batches, reweighted to the 1,685 companies in the study base. An automated classifier was built first and rejected. Its numbers are further down this page, because they are the most useful thing on it.
This page sets out the categories, the test that governs the sort, the mechanical rules that resolve most cases without judgment, what the hand classification found, and what none of it can support.
Which figures come from these rules and which do not
Every figure that requires knowing whether a person wrote a specific email comes from the 190 hand verdicts, reweighted. That is the 93.4% never-answered rate and its 6.6% complement, the marketing send bound of 8% to 26%, the observation that none of the sixteen fast replies we hand checked was written by a person, and the comparison of median lag between human and automated replies.
Figures that need no such judgment do not come from these rules and are not claimed to. Arrival times, the 545 replies inside five minutes, the 499 arriving outside working hours, word counts, question counts, call to action counts, link counts, form attributes and chat installation all come from the coded corpus of 2,102 readable emails and from the study workbook. They are mechanical, and they are labeled that way on the demo response study method page.
Not every figure in the set comes from those 190. The split above is what each figure actually rests on.
The three categories
Every email received after a demo or contact sales request is placed in exactly one of three categories.
HUMAN. A person read this specific request and wrote something a template could not have produced. The email refers to something particular about the request, the company that submitted it, or an action the sender has taken.
ACKNOWLEDGMENT. Automated. The category includes a named sales representative sending a generic first touch, calendar link invitations, and any variation of "we have received your request". A sales representative's name at the bottom of a sequence step does not move an email out of this category.
MARKETING BLAST. Not a reply at all. Newsletter, webinar invitation, content download, product announcement. The submitted address went onto a list. These emails answer nothing that was asked.
The governing test
One question decides the category.
Could this exact email have been sent, unchanged, to any other person who filled in that form?
If yes, it is not a human reply. The test is deliberately strict, and it is strict in the direction of undercounting human replies rather than overcounting them.
Merge fields prove nothing
First name, company name and email domain all came from the form. An email that uses them is reproducing data the buyer supplied, which is what a mail merge does.
This is not a stylistic judgment. Every submission in the study came from one identity, a VP of Engineering at a 200 person software company in San Francisco, and no form in the study offered a free text message field. There was no information for a sender to respond to beyond the fields we typed in, with two exceptions that cut the other way: the identity carried a LinkedIn profile and the company had a website. A sender who went looking could find something on either that the form had not supplied, and a reply that reports doing exactly that is the clearest instance of the feature this rulebook is built to detect. Those two are where the replies that passed the test had been: one named the category the company sells into, another named something off the LinkedIn profile. Personalization drawn entirely from those fields is therefore reproducible by software with no reader involved, and it is not evidence that a person was involved.
This rule removes the largest single source of misclassification. Automated first touches read as personal precisely because they are built from the form.
The working hours rule
An email arriving before 08:00, from 18:00, or at a weekend in the sender's own local time is treated as automated.
Sender local time is recovered from the timezone offset carried in each email header, not from the recipient's clock and not from an assumption about where the company is based. An email sent at 03:00 recipient time by a sender eight hours ahead is an 11:00 email and falls inside working hours.
The rule was checked against hand verdicts on a first batch of 36 out-of-hours emails and held 34 times. Both exceptions described something the sender had actually done. One had telephoned and left a voicemail. The other explained that demonstrations were only run in a language other than English. In both cases the sender was reporting an action, which is the signature the rule is designed to catch, and both were classified HUMAN despite the arrival time.
The bulk mail infrastructure rule
An unsubscribe link, a preferences link, or a role sender address such as info@ or noreply@ means automated.
The reasoning is mechanical rather than stylistic. A person replying from a mailbox does not attach an unsubscribe footer, because that footer is inserted by sending infrastructure built for lists. Across 30 hand classified emails carrying that infrastructure, none was judged human.
The hand classification
190 emails were read and judged one at a time: 17 human, 162 acknowledgment, 11 marketing blast. The 8.9% human share of that sample is not a population rate, because the strata were sampled at deliberately different fractions.
Sampling was stratified across the 2,102 readable emails in the usable base, so that verdicts reweight to all 1,685 companies. Stratification used the structural features above rather than content, so the strata could be defined before any email was read.
| Stratum | Population emails | Hand verdicts | Human |
|---|---|---|---|
| B, bulk mail infrastructure or role sender | 1,060 | 30 | 0, 0.0% |
| C, personal sender, no direct ask | 186 | 35 | 2, 5.7% |
| D, personal sender, makes a direct ask | 856 | 125 | 15, 12.0% |
Stratum D is the deciding one, because it holds the emails that could plausibly go either way and it carries most of the weight in the reweighted figure. It was drawn twice, independently. The first draw returned 8 human of 65, 12.3%. The second returned 7 of 60, 11.7%. Two independent draws landing 0.6 points apart, on a stratum whose rate differs from stratum B by twelve points, is the condition under which the stratification is doing useful work.
The reweighting is stated so it can be rerun. For each of the 1,685 companies, the probability that no human replied is the product across strata of one minus the stratum rate, raised to the number of that company's emails in that stratum. Summed across the 1,685 that gives 1,573.28 companies, 93.4%. A 20,000 draw parametric bootstrap over the three stratum rates returns 90.3% to 96.1%.
The 6.6% complement, about 112 companies, is the same modeled expectation read the other way. It is not 112 identifiable companies and it cannot be turned into a list.
What separated a human reply from an acknowledgment
One thing, and nothing else came close: the sender reporting an act they had personally performed on this specific submission, and naming the fact they found. Present in 14 of 17 human emails and 5 of 162 acknowledgments.
Asking a question separated nothing, 71% in both classes. "Thank you for your interest" appeared in 18% of human emails and 17% of acknowledgments. A named personal sender address appeared in 17 of 17 human emails and 152 of 162 acknowledgments, which makes it worthless as a signal.
The classifier that was built and rejected
An automated classifier was built first. It weighted four features: the interest phrase, the demo request phrase, a named sender address, and speed of reply.
Three of those four appear at the same rate in both classes. The fourth runs backwards, because among the 190 hand labeled emails the median lag was 765 minutes for the ones a person wrote and 281 minutes for automated acknowledgments.
Measured against the 190 hand verdicts:
| Measure | Value |
|---|---|
| Emails scored | 190 |
| Agreements with the hand verdict | 70, 36.8% |
| Emails the classifier called human | 116 |
| Emails the hand verdicts call human | 17 |
| Of the 116, actually human | 11 |
| Precision on the human class | 9% |
| Of the 17 human emails, found | 11, 65% |
A classifier that calls 116 emails human when 17 are is not a classifier that needs tuning. Nine times out of ten, when it said a person had written the email, a person had not.
Measured against the first batch of 36 alone, the same classifier agreed 10 times and called 16 emails human where the correct answer was 2. The failure was visible at 36 and confirmed at 190.
The method was abandoned rather than tuned, because tuning would have produced a classifier fitted to the batch it was tuned on and the fault was in the premise. The premise was that the words in an email indicate whether a person wrote it. They do not. Autoresponder phrasing such as "thank you for your interest" and "your demo request" refers to the specific request, which is exactly why templates contain it.
Nothing in the published set rests on that classifier. Every human reply figure comes from the hand classification described above.
What these rules cannot support
They cannot resolve differences between site types. The 190 verdicts were allocated to measure one overall rate precisely. Split across subgroups they give intervals 6.5 to 10.6 percentage points wide that overlap heavily, eleven of the eighteen cells contain no human verdicts at all, and the stratum B long-form cell holds zero hand labels against 124 population emails. Settling the chat installed against no chat comparison would take roughly 880 further hand classifications. The form length comparison cannot be settled at any sample size, because only 73 long-form emails exist in stratum D. No subgroup never-answered rate is published from these verdicts.
They hold no outcome data of any kind. Classification records what an email was, not what happened next. Nothing here records a meeting, a deal, a renewal or a lost buyer.
They undercount rather than overcount. A person writing a short generic reply from a company mailbox with a standard footer is classified as an acknowledgment. That error is accepted, because the opposite error, counting sequence steps as human replies, would inflate every headline figure in the research.
They prove absence, not presence. The mechanical rules can establish that an email was not personally written. They cannot establish that one was. That is why 800 emails surviving the exclusions is not a count of human replies, and why the human reply figure needs hand classification at all.
They apply to email only. They say nothing about telephone responses, which the underlying dataset does not capture.
Human class proportions are directional. Every proportion computed inside the 17 human verdicts moves by about six points if one email is reclassified. They are published as direction and read that way.
How to cite this
GTM Clarity, How we decide whether a human replied, 2026. Demo response data collected 2023. Research directed by Terry Wilson, fieldwork by the GTM Clarity research team.
The figures these rules produce are reported in The Buyer Wait Time Report 2026.
Terry Wilson is the founder of GTM Clarity and CEO of ChatMetrics, which has delivered over $5 billion in qualified pipeline and 300,000+ leads for B2B clients across SaaS, services, and industrial sectors. Before founding ChatMetrics, Terry was National Sales & Marketing Manager for a $1B enterprise, leading more than 350 people across Australia. He built GTM Clarity's AI on a corpus of 3M+ real B2B sales conversations that delivered $5B+ pipeline across 200+ companies.
Keep reading
The Buyer Wait Time Report 2026
Think about what has to happen before a person fills in your demo form.
Read →Statistics62 B2B buyer response statistics, from five years of original research
Base for this section: 4,265 B2B technology websites, each above 10,000 monthly visitors, classified in 2021, again in 2025 and again in 2026 at the same addresses. The 2021 and 2025 figures are observed states with no coder judgment about intent, no estimation and no sampling. The 2026 wave sampled the three big groups and tried every site in the three small ones, so its figures carry confidence intervals and the first two waves do not. It is also a different instrument: 2021 and 2025 classified a site by opening the widget on one visit, while 2026 held a conversation and waited three minutes for a person.
Read →MethodThe demo response study, 2023, method and dataset
93.4% of 1,685 B2B companies never put a human in front of a buyer who had asked to see the product, 1,573 of 1,685, 95% confidence interval 90.3% to 96.1%.
Read →