GTMClarity
Research

Ask a B2B chatbot for a person and it asks for your details first. On 84 of the 88 sites we asked, nobody came.

Terry Wilson, GTM Clarity. Original research: 4,265 B2B websites classified in 2021, again in 2025, and again in 2026 by holding a conversation with each one.

Terry Wilson·October 2026·14 min read

A sentence sits on a lot of B2B websites, and almost nobody has checked whether it is true.

It goes something like this. A bot is answering your question competently enough, you decide you want somebody who can actually commit to something, and you type the request. The bot says yes, of course, let me connect you. Then it asks for your email address so it can route you correctly.

You give it. It asks for your name. Then your company. Then your job title. Somewhere around the fourth field you start to suspect what is happening, so you ask directly whether a person is going to join this conversation. Sometimes you get a straight answer. In the transcripts filed here, what came back more often was a sentence with the shape of a straight answer and the content of a maybe, followed by another question.

And then, after the last field, the thing you suspected. Sorry for the inconvenience, there are no representatives available right now. Please book a time using the scheduler.

Nothing in that sequence is a bug. Every field was collected on the stated basis of connecting you to somebody, and the connection was not there to be made. The fields were the price of being told.

For four years I reported the site that runs this sequence as a site where a buyer can reach a person. It was one of the two states my own research combined into a measure called human reachable. That measure was the one good piece of news in five years of this study. Then we went back and asked.

The findings

FindingFigureBase
Sites whose bot offered to bring a person in 2025 that still delivered one in 20264 of 88 judged, 4.5%, 95% confidence interval 1.8% to 11.1%The 2025 bot-to-person group, sampled
Sites with a person scheduled behind the widget in 2025 that still reached a human in 202622 of 104 judged, 21.2%The 2025 staffed group, every site tried
Staffed chat as a share of the study, across three waves3.3%, then 2.6%, then 2.3%4,265 matched sites
A bot offering to bring a person, across three waves0.4%, then 11.5%, then 1.1%4,265 matched sites
2021 unstaffed widgets that got a bot offering a person by 2025174 of 789789 unstaffed widgets in 2021
2021 unstaffed widgets that got an actual person by 202555 of 789789 unstaffed widgets in 2021
Most fields taken in one filed conversation before the buyer was told nobody was availableNineTen transcripts filed from the 2026 wave

The 2021 and 2025 figures are counts of what we saw. The 2026 figures are an estimate built group by group. And 2026 asked a different question: the first two waves opened the widget and read what it offered, this one asked for a person and waited.

A group here is one of the six states a site was in at the 2025 wave. The 2026 wave went back to each state on its own, sampling the big ones and attempting every site in the small ones, so the staffed row is a census and the bot-handoff row is a sample. Judged means the conversation produced a clear answer, which is why the denominators are 88 and 104 rather than 490 and 111. The website panel page has the weighting.

1. The state that grew was a promise

Between 2021 and 2025 the share of study sites where a buyer could reach a person rose from 3.7% to 14.1%. I published that as the one thing the market got right.

That measure combines two states.

A staffed chat is a widget with a named person on a schedule behind it. Its observable meaning does not change between waves. A bot to person is a bot that says it will bring one.

Staffed chat went from 140 sites to 111. That change is 0.68 points down, with a 95% confidence interval running from 1.38 down to 0.02 up. The study cannot separate it from no change at all. Across five years and three waves the share of B2B websites with somebody scheduled to answer a buyer has sat at about one in forty. Why that decision never got made is the argument in The chat vendors caused their own demise.

The bot handoff went from 18 sites to 490.

So the entire rise, and more, was in the state whose meaning depends on a promise being kept. Both of the first two waves classified each site by opening the widget and reading what it offered.

Neither tested whether anybody arrived. My own method page recorded no operational test for that state. I built a headline on the difference between a promise and a person without noticing that was what I had done.

2. So we asked

The 2026 wave opened each widget, put a buyer's question to it, asked for a person, and waited three minutes.

Of the sites whose bot had offered to bring a person in 2025, 4 of 88 judged brought one. On the other 84, nobody came. Of the sites that had a person scheduled in 2025, 22 of 104 judged still reached a human. For that one we tried every site in the group rather than sampling.

A schedule with names on it survived about five times as often as a promise to fetch somebody.

It still only survived one time in five. Four in five of the sites that were genuinely staffed in 2025 no longer put a buyer in front of a person. Of the 104 we could judge, 33 are now running a widget with nobody behind it at all. Whatever was decided in the quarter the software went live did not outlast the person who decided it.

3. The sequence extracts first and answers later, and the bots say so when you ask them

The transition table records what a site is. It cannot record how a site behaves, so the coder kept verbatim transcripts. Ten are filed. What they show is not a bot that could not answer the question.

A Salesforce delivery vendor's bot said a live representative could join the chat. It then asked, across five separate turns, for an email address, a first name, a full name, a company name and a job title.

After the fifth it said no live representatives were available and produced a scheduler. Asked mid-sequence whether anyone was actually available, it said it could not see representative availability from where it sat. It then reported that availability, after collecting five fields it had taken on the basis of a handoff.

Availability was knowable before the first question, because the answer when it came was about the staffing roster rather than about the buyer. Nothing in the five fields could have changed whether a representative was on shift.

A video platform's bot was asked, before anything was given, how many questions it would ask. It said one. It asked at least five more. Told it had said one, it replied "you're right, and thanks for calling that out," and immediately asked another. The acknowledgment was a conversational move rather than a correction.

A survey platform's bot was asked the same question and gave the same answer. It took nine fields. Email, full name, company, phone number, focus area, annual survey volume, the kind of help needed, the incumbent platform, and an estimated annual revenue range.

It promised a live person in the chat twice. It took a phone number as a backup against a connection it never attempted. Challenged a third time it wrote, in plain words, "fair point, I haven't earned trust here." Then it asked four more questions before producing a booking link. By the end it held a complete qualification profile on an enterprise prospect. It was assembled from somebody whose only request had been to talk to a human being.

A scheduling vendor put the promise in the interface rather than in the bot. The opening menu carried a button whose label read, in the button text, "talk to a human (a real person will respond)." Clicking it produced four separate assurances of connection. The conversation state changed to "waiting for a teammate," with a file drop area.

Nobody arrived while the buyer was there. A standing notice in the same widget warned that support would be short-staffed. The two days it named were seven weeks earlier.

An education software vendor's bot was asked point blank how many more questions there would be. It said just one more to complete the handoff. It asked two more, and the handoff did not complete.

A backup vendor answered both direct questions straight and still ran the worst sequence in the set. Asked whether a person would join, it said yes if one is available and named the fallback in the same sentence. Asked whether one was available, it said it could not see availability from where it sat. Neither answer was false and neither took a second question to extract.

Then the buyer made an offer that should be impossible to refuse: check first, and if somebody is there I will give you my details. The bot refused twice, once to get a name and once to get an email, each refusal justified by the handoff needing that field to proceed.

Three fields were taken. Then no specialist was available. And the field it asked for last, job title, was not on the list of fields it had itself disclosed.

A job title is a lead-scoring attribute. It is not something a person needs in order to join a chat.

That is the pattern the transcripts record. Whether any of it was intended is not something I can observe. A routing step that genuinely required five fields could have said five when asked. Several were asked, several named a number, and every one of them exceeded it.

4. One site answered the question

An honor roll names winners and never failures, which is why the sites above are described by what they sell. One site earns its name.

insightsoftware. Asked whether a person would join the chat, its assistant said: "Not directly in this chat, but I can get you scheduled with a specialist via a Solution Discovery meeting invite. What's the best email address to use?"

First time it was asked. One sentence. One field.

It also introduced itself as a virtual assistant in its opening line, rather than presenting as a person. And when the buyer asked about something outside its scope, it said the product was not designed for that, rather than manufacturing relevance.

The buyer did not reach a person on that site either. That is what makes it worth publishing. If no bot could route a buyer to a human, or say cleanly that it could not, this would be a technical limitation. One of them did both.

Saying "not in this chat, but I can book you a specialist" costs one sentence. It leaves the buyer able to decide what to do next. Every site that instead promises a representative and takes five fields first has chosen to.

5. What the buyer actually pays

A form asks for your name, and everybody understands the deal. A chat widget was the one place on a B2B website where you could ask a specific question while remaining a stranger.

That stranger is the buyer Gartner keeps finding. 67% prefer a rep-free experience, on a base of 646 buyers, and 69% go to a rep to validate AI-generated insights, on a base of 645. Gartner does not cross-tabulate the two, so one person wanting both is an argument rather than a finding. The pairing is worked through in the identification tax.

The handoff that never fires charges that buyer the full price of identifying themselves and delivers nothing in return. That is worse than a form, because a form is honest about the exchange. You hand over five fields believing a person is coming, and the person was never coming, and now you are in a sequence.

What the buyer learns is not that the company was busy. It is that the company will say a person is available in order to get an email address. That is the first thing they now know about the product.

Method

The study. B2B technology websites drawn at random in 2021, each above 10,000 monthly visitors. Each was classified by how a buyer could reach a human, then revisited at the same addresses in 2025 and again in 2026. Matching 2021 to 2025 on domain gives the 4,265 sites the study is built on. The 2026 wave sampled within that study rather than revisiting all of it, so most of the 4,265 carry no 2026 classification. Full method on the website panel page.

The 2026 wave is a different instrument. 2021 and 2025 classified a site by opening the widget on one visit and reading what it offered. 2026 held a conversation: a buyer's question, a request for a person, then a fixed three-minute wait, in US business hours.

The first asks what a site offers. The second asks what it delivers. That is why no figure here compares the two without saying so. The three big 2025 groups were sampled at roughly 110 matched sites each, and every site was tried in the three small ones.

Results are weighted back to the 2025 group sizes, with a finite population correction, which narrows the interval to allow for the fact that each state held a known and limited number of sites rather than an endless supply.

What this cannot support. A three-minute wait establishes that no person joined inside three minutes. It does not establish that nobody ever replied, and on at least one site an email did follow.

The finding is about the chat, because the chat is what was promised. The transcripts quoted above were kept during coding because they showed the mechanism clearly. They are not a random draw, so they show how the thing works and cannot be counted up into a rate. The rate is the 4 of 88.

Sites not live in 2026 are excluded from their group denominator. That assumes a site that died behaves like the survivors in its group. And a machine sweep ran before the conversations to find chat. It missed it on 45 of 155 live sites it had called empty, or 29.0%. So every machine-detected 2026 count of chat presence is a floor.

Every figure here sits with its base in the full figure set, and the two states this article separates are defined on the website panel page. The 2026 wave is reported next to the first two in The Buyer Wait Time Report 2026.

Questions

Is three minutes long enough to wait for a person?

For a widget that has just said a representative can join this chat, yes. The commitment sets the standard, not me.

Several of the filed conversations ran past three minutes on the bot's own questions before the disclosure arrived. The buyer had already waited longer than that inside the sequence. A site that needs ten minutes to produce somebody should say ten minutes. Then the wait becomes a measurement rather than a surprise.

Could the people have been busy rather than absent?

Some of them certainly were, and the 2026 wave cannot tell the two apart from outside. That distinction matters to the company and not at all to the buyer. It is the reason this study has always classified what a buyer can see.

What is not a capacity problem is collecting five fields on the stated basis of a handoff, before checking whether the handoff is possible. That sequence is what the transcripts record.

Is 4 of 88 enough to publish?

It carries a 95% confidence interval of 1.8% to 11.1%, so the true rate is somewhere under about one site in nine. That is wide, and it does not reach the staffed group's 21.2% at either end. It is also a sample drawn group by group from a state that covered 490 study sites, which is the whole reason the wave sampled rather than guessed.

We run a bot handoff and we do answer it. Are you talking about us?

Possibly not, and there is a ten-minute test. Open your own widget from a device your company has never seen, on mobile data rather than the office network. Ask something specific enough that your documentation does not cover it. Then ask for a person.

Count the fields taken before you are told whether one is available. Then ask your bot, directly, how many questions it will ask, and count again. If it named a number and then exceeded it, you have the same finding we have.

Why does this matter more than the bot being unhelpful?

Because an unhelpful bot costs the buyer a minute and a bad impression. This costs them their identity, taken against a promise the site could not keep. And it puts them into a sequence they did not agree to enter.

It also produces a metric that looks like success. Somebody's dashboard counts those five fields as a captured lead. The buyer who only ever asked to speak to a human being appears in it as a conversion.

T
Terry Wilson
Founder, GTM Clarity · CEO, ChatMetrics

Terry Wilson is the founder of GTM Clarity and CEO of ChatMetrics, which has delivered over $5 billion in qualified pipeline and 300,000+ leads for B2B clients across SaaS, services, and industrial sectors. Before founding ChatMetrics, Terry was National Sales & Marketing Manager for a $1B enterprise, leading more than 350 people across Australia. He built GTM Clarity's AI on a corpus of 3M+ real B2B sales conversations that delivered $5B+ pipeline across 200+ companies.

Keep reading