GTMClarity
Research

Five years of AI hype, and B2B websites went backwards

Terry Wilson, GTM Clarity. Original research, 4,265 matched B2B websites classified in 2021, again in 2025, and again in 2026.

Terry Wilson·October 2026·19 min read

If you have a question about a B2B product and you would like it answered before you hand over your name, your options have gotten worse over the last five years. Not better. Worse. That is the opposite of what five years of language-model progress was supposed to do, and it is measurable.

Everyone has seen the little chat window in the bottom corner of a B2B website. Over the last five years you would have every reason to assume there are more of them than there used to be. Every market forecast says so. Every vendor deck says so. The whole category has spent four years selling the idea that buyers now expect to talk to something the moment they land.

We tracked the same 4,265 B2B websites from 2021 to 2026. There are fewer.

I have been classifying these sites since 2021, one visit per site per wave, recording what a buyer arriving on the page could actually do. Not what the vendor claimed, not what the forecast projected. What was there.

Those years produced the largest step change in language technology in the history of the field. Over the same years, the share of B2B websites offering a buyer no conversational path of any kind went up. More sites turned chat off than turned it on, and it happened while the category was reporting its best years.

Both things can be true at once. Category reports count vendor revenue and seats sold, which are supplier-side numbers. A vendor can raise prices, move upmarket and grow revenue while appearing on fewer websites.

There is a second story inside that figure. For four years I read it as better news than the headline suggested. The third wave changed my mind, and the reason is that we stopped counting widgets and started asking them for a person.

The findings

Finding202120252026
No chat at all64.0% (2,731 sites)72.3% (3,083 sites)74.1%
Staffed chat, a person scheduled behind it3.3% (140 sites)2.6% (111 sites)2.3%
Bot to person, a bot offering to bring one0.4% (18 sites)11.5% (490 sites)1.1%
Unstaffed widget18.5% (789 sites)1.6% (69 sites)7.0%
Bot only13.8% (587 sites)12.0% (512 sites)15.6%
Sites removing chat against sites adding it, 2021 to 2025753 removed401 added
2025 bot handoffs that still delivered a person in 20264 of 88 judged
Sites staffed in 2025 that still reached a human in 202622 of 104 judged

Every row sits on the matched panel of 4,265 B2B websites classified in all three waves. The 2021 and 2025 figures are counts of observed states. The 2026 figures are an estimate built group by group and carry confidence intervals. Those are in the method and bounds section. The 2026 wave is also a different instrument. It held a conversation rather than looking at a widget.

The 2026 column is an estimate rather than a count, because the third wave went back to each 2025 state separately instead of revisiting all 4,265 sites. Judged means the conversation gave a clear answer, so those denominators are the sites we could read. The sampling is set out on the website panel page.

More doors closed than opened

No chat at all rose from 64.0% to 72.3% between 2021 and 2025, 2,731 sites to 3,083, and stood at 74.1% in 2026.

753 websites turned chat off between the first two waves. 401 turned it on. Same sites, four years apart. The 2026 reading carries a 95% confidence interval of 69.1% to 79.1%, which does not settle whether anything moved after 2025.

These are the same websites in both waves, so a site that removed chat is observed removing chat rather than inferred from a shifting sample. That is the whole reason the number is worth anything. The full panel method sets out how the matching was done, for anyone who wants to check it.

A buyer landing on one of those sites in 2025 meets a finished, confident page with nowhere to put a question. Four years earlier, the same page had somewhere.

What died, and then came back

The unstaffed widget fell from 18.5% to 1.6%, 789 sites to 69. In 2026 it was back to 7.0%.

That fall was the single largest movement between the first two waves. Here is where those 789 sites went.

Outcome in 2025SitesShare of 789
Switched it off39149.6%
Put a person behind it557.0%
Put a bot in front of it, offering a person17422.1%
Put a bot in front of it, offering nothing13817.5%
Still unstaffed313.9%

Removal beat putting a person behind it by seven to one. Faced with a widget nobody was answering, one company in fourteen staffed it, one in two deleted it, most of the rest handed it to software, and 31 left it sitting there unstaffed. The third option, a bot that offers to fetch somebody, was taken three times as often as hiring anybody to answer.

And in 2026 the widget with nobody behind it returned, to 7.0% of the study with a 95% confidence interval of 3.8% to 10.2%. The interval excludes the 2025 figure, so that one genuinely moved. Of the 104 staffed sites we could judge in 2026, 33 are now running a widget with nobody behind it. Whatever the first two waves looked like, this was not a habit the industry broke.

The unstaffed widget was always a strange object. It sat in the corner of the page performing availability and delivered messages into a queue nobody was watching.

It was not free. These were paid products on contracts, and companies went on paying while nobody answered, which is clearest in what happened to Drift's customers. What they were was cheap next to the thing that would have made them work. A license is a line item, signed off once. A person on a rota is headcount, argued for every year, usually by somebody else.

That asymmetry is why removal was easy. Taking the widget off cost nothing operationally, because nothing had been built around it. No rota to unwind, no promise to the buyer to withdraw, no process to retire.

So the install base measured what a marketing budget would carry, not what anyone had committed to. Nearly a fifth of the B2B web was running one in 2021, and that was never a measure of conversational maturity.

Then it died the way line items die. Somebody goes down a list looking for things to stop paying for, and a widget nobody answers is the easiest line on the page. The cost sits on an invoice. The return sits nowhere. On the evidence in front of them that reading was fair, because a widget nobody answers is not returning anything.

What a cost exercise cannot price is the thing being given up, and here it had never been in place long enough to measure. Not a channel. The ability to answer a buyer's question at the moment they asked it. That was never on the invoice, so it was never in the decision.

The companies that removed it were not shutting down a working channel. They were tidying up after a decision made years earlier by somebody who had probably left.

The thing that grew was a promise

Staffed chat went from 140 sites to 111. The bot that offers to bring a person went from 18 sites to 490.

For four years I published those two states as one number, called human reachable. It rose from 3.7% of the study to 14.1%. I read that as the expensive tier growing while the cheap tier collapsed.

They are not one thing. A staffed chat is a widget with a named person on a schedule behind it. A bot to person is a bot that says it will bring one. Both of the first two waves classified a site by opening the widget and reading what it offered, which records the offer. Neither tested whether anybody arrived.

So the state that grew is the one whose meaning depends on a promise being kept, and the state whose meaning is fixed has not moved.

Staffed chat was 3.3% of the study in 2021, 2.6% in 2025 and 2.3% in 2026. The 2021 to 2025 change is 0.68 points down, with a 95% confidence interval of 1.38 down to 0.02 up. The study cannot separate that from no change at all.

Across five years and three waves, the share of B2B websites with a person scheduled to answer a buyer has stayed at about one in forty.

In 2026 we tested the promise. The third wave opened each widget, asked a buyer's question and waited three minutes.

Of the sites whose bot had offered to bring a person in 2025, 4 of 88 judged still delivered one, 4.5% on a 95% confidence interval of 1.8% to 11.1%. Of the sites that were genuinely staffed in 2025, 22 of 104 judged still reached a human. That is 21.2%, and we tried every site in that group. What those bots said when they were asked for a person is in Ask a B2B chatbot for a person and it asks for your details first.

A schedule with names on it survived about five times as often as a promise to fetch somebody. It still only survived one time in five.

The buyer's side of that arithmetic is plainer than any of it. Thirty-nine times in forty, the question you arrived with has nobody scheduled to answer it, and that has been true the whole time.

The demo response study I ran alongside this one puts a floor under how low the base was. It was fielded in 2023, so treat it as a baseline rather than a current reading.

Of 1,685 B2B companies that received a real demo request through their own website form, 93.4% never put a human in front of the buyer. A website offering no route to a person is consistent with what those companies do when a buyer takes the route they do offer.

Business buyers wait where consumers do not

Staffed chat went from 3.11% to 2.54% across 2,795 B2B sites, and from 3.80% to 2.85% across 1,368 consumer-facing sites. The bot handoff went from 0.39% to 12.81% across the B2B group and from 0.51% to 9.58% across the consumer-facing group. 102 study sites carry both sector flags or neither and are excluded.

I published this comparison for four years on the combined measure, where consumer-facing sites led by about five points in 2025. Two things are wrong with that.

The bases it was published on cannot be reproduced from the source workbook by any reading of its sector columns. That version of the comparison has no surviving derivation. And on the bases that can be reproduced, the combined measure puts the B2B group ahead in 2025, 15.35% against 12.43%. The B2B group carries more of the bot handoff.

What survives the split is much smaller than what it replaces. Look at staffed chat, the one state whose meaning does not change between waves. Consumer-facing sites are ahead in both waves, by less than a point, and both groups drift down together. On the combined measure the lead changes hands, which is why I publish no significance test on it. A comparison that changes sign depending on which of the workbook's sector columns defines the groups is not a finding.

The usual explanation for a consumer lead is cultural. Consumer businesses are closer to their customers, more service-minded, less enterprise-brained. I do not believe it, and if there is a mechanism it is simpler than that.

A consumer-facing team can staff chat on Monday and read the conversion difference by Wednesday. The decision to buy sits on the same page as the conversation, the cycle is minutes long, and the experiment settles itself. Nobody has to win an argument, because the data arrives before the next meeting.

B2B has a long sales cycle and no attribution model that anybody in the room actually trusts. A conversation in March cannot be cleanly connected to a closed deal many months later, and everyone involved knows it. So the question never gets settled by evidence. It gets settled by whoever is most confident in the meeting.

That is a structurally biased contest, and the bias runs one way. The cost of staffing chat is visible, immediate, and appears on a line item with a name. The benefit is delayed, diffuse, and shows up inside a number that four other teams also claim.

When people argue about a cost they can see against a benefit they cannot, the cost wins. It wins in most quarters, at most companies, regardless of which side is right.

I cannot prove any of that from this study, and on these bases the gap it would explain is under a point. What the split does show clearly is that the feedback loop argument does not need a sector gap to do its work. Neither group solved this. The state that costs money to hold stayed where it was in both.

The one channel that could have answered without asking who you were

This retreat happened at a particular moment, and the timing is the most damning part of it.

A form takes your name. 70.7% of the demo forms we tested in 2023 take your phone number as well. A phone call identifies you by definition. A chat widget does not. You can type a question into a box and get an answer while remaining a stranger.

So a company removing chat is not removing one channel among several. It is removing the only one that did not charge the buyer for the privilege of asking something. What is left is a page that will answer you once you have handed over a name, a company and a number. That is a lot to pay for a question you are not yet sure you want to ask.

Now put the timing against it. Across one fieldwork window in late 2025, Gartner found 67% of B2B buyers prefer a rep-free experience, on a base of 646, and 69% go to a rep to validate AI-generated insights, on a base of 645. Read as the same person, which is inference rather than a cross-tabulation, that is a buyer who needs one question answered late and without identifying themselves. I take the pairing apart in the identification tax.

B2B spent those four years removing the only mechanism on the page that served that buyer.

Here is my guess at the loop, and it is inference rather than measurement. Research moved to places a company cannot see, so attribution got harder. Anonymous engagement produces no attributable revenue on a dashboard, so it looked like a cost with no return. It got cut.

Cutting it pushed more of the research off-property, which made attribution harder still. Removal beat staffing, and the thing being removed was the last visible point of contact with a buying process that had gone dark.

The script was probably never the problem

Bot only moved from 13.8% to 12.0%, 587 sites to 512, and was 15.6% in 2026.

Down 1.8 points across the first two waves, then up 3.6 points in the third. That third reading carries a 95% confidence interval of 12.1% to 19.0%, which does not rule out no change in the third step.

Look at what those four years contained. Language models went from a curiosity to a line item in every software budget on earth, and every chat vendor rebuilt its product around them. A company that switched its bot off in 2021 because the thing was useless has had four years of increasingly good reasons to switch it back on.

If better automated conversation were the fix, that share should be climbing steadily. It has drifted down and then back up, inside intervals that cannot separate either step from no change.

My guess at why, and I cannot prove it from this data, is that the companies removing chat were not removing a bad conversation. They were removing one that nobody was ever going to finish. A buyer asks something real: pricing, or a security review, or whether the thing integrates with their stack. They get an answer that is competent and generic.

Then they reach the point where the answer needs somebody with authority. In 2021, on 89.7% of every site that had chat at all, a buyer could not reach a human through it. By 2025, on most sites, there is nothing to hand off from at all. Making the first half of a conversation better does very little when there is no second half.

There is a second reason, and it is the one that removes chat's whole advantage. A widget that asks for a work email before it will answer anything has stopped being a different channel. It is a form in a smaller box.

The buyer came to the corner of the page precisely because it looked like the one place they could ask something without identifying themselves first. The moment it charges the same fare as the contact page, there is no reason to prefer it, and no loss in taking it away.

That is the part the buyer is left with. Not the quality of the bot, which this study never measured, but the sentence where the conversation stopped being able to help and nobody arrived to take it further.

The 55 sites that put a person behind their unstaffed widget were the only ones that addressed what was missing. They were outnumbered seven to one by the ones that deleted it. Deleting it is the right call if you were never going to answer it. They were also outnumbered three to one by the ones that put a bot in front of it instead.

An unanswered invitation is worse than no invitation. A bot promising a person who never arrives is worse than both.

Method and bounds

B2B technology websites were drawn at random in 2021, each above 10,000 monthly visitors. Each was classified by how a buyer could reach a human, then revisited and reclassified on the same scheme in 2025. The 2021 wave holds 5,007 records, 5,004 of them carrying a complete single classification, across 4,832 distinct domains.

Matching the waves on domain gives the 4,265 sites that form the base here. Full method on the website panel page, and every figure here is in the full figure set. The five-year picture across both studies is in The Buyer Wait Time Report 2026.

Tests and intervals on the 2021 to 2025 change. Because the same sites appear in both waves, the change is tested with McNemar's test, the standard test for a before-and-after shift in the same subjects, and the interval is a Wald interval on the paired difference. A low p-value means a change that size would rarely appear if nothing had really moved; an interval that straddles zero means the study cannot tell the change from no change.

No chat at all, p below 0.0001, 95% CI +6.7 to +9.8pp. Staffed chat, p=0.064, 95% CI -1.4 to +0.02pp, which is why no fall is asserted. Bot to person, p below 0.0001, 95% CI +10.1 to +12.0pp.

Any unstaffed widget, p below 0.0001, 95% CI -18.1 to -15.7pp. Bot only, p=0.009, 95% CI -3.1 to -0.5pp.

The 2026 wave. Coded 16 to 18 September 2026 by opening the chat, asking a buyer's question and waiting a fixed three minutes, in US business hours. The three big 2025 groups were sampled at roughly 110 sites each, and every site tried in the three small ones.

Results are weighted back to the 2025 group sizes, with a finite population correction, which narrows the interval to allow for the fact that each state held a known and limited number of sites rather than an endless supply. 95% intervals: no chat 69.1% to 79.1%, bot only 12.1% to 19.0%, any unstaffed widget 3.8% to 10.2%, staffed chat 0.4% to 4.1%, bot to person minus 0.7% to 2.9%.

A machine sweep ran first to find chat. It missed it on 45 of 155 live sites it had called empty, or 29.0%. So every machine-detected 2026 presence count is a floor. Sites no longer live are excluded from their group denominator, and 50 of those 93 domains redirect to a different company.

Segment split. Groups are defined by the 2025 workbook's own B2B and B2C columns, read strictly. A site marked in exactly one is assigned to it. A site marked in both or neither is excluded.

That gives 2,795 B2B sites, 1,368 consumer-facing sites and 102 excluded. No significance test is published on the comparison. An earlier version of this article used two different segment bases. They cannot be reproduced from the workbook, and they are withdrawn along with the figures that sat on them.

Demo response study, 2023 fieldwork, a baseline rather than a current reading. The 93.4% counts 1,573 companies out of 1,685, and its 95% confidence interval runs from 90.3% to 96.1%. The form composition figure, 70.7%, counts 1,192 forms out of the same 1,685.

Gartner. The two releases rest on one fieldwork window, August to September 2025, and were published in March 2026 and May 2026. They state two bases, 646 buyers and 645, and Gartner publishes no cross-tabulation of the groups, so reading them as the same buyers is inference.

Two limits decide how far this travels. Each site was seen once per wave. So a company staffed only during business hours in one timezone is recorded as whatever it was at the moment of the visit. And this is what a buyer could see from outside a website, which says nothing about what happens after they engage.

FAQ

Chat moved into the product. Doesn't that explain the drop on marketing sites?

Partly, and the study cannot see inside a logged-in product, so I cannot size it. It does not explain the whole shape, and it does not touch the people this piece is about.

In-product chat serves customers. The 352 net removals are front doors. The person standing at a front door has not signed up, has no login, and cannot follow a conversation into a product they do not have.

They are the prospect who never becomes a lead, never reaches the CRM, and never appears in the pipeline as anything at all. Moving support into the product is a good decision that answers a different question from the one this data raises.

Category reports show this market growing. Are they wrong?

Probably not wrong, just measuring a different thing. They count vendor revenue and seats sold, which is supplier-side. This counts websites, which is buyer-side. A vendor that raises prices, kills its free tier and moves upmarket grows revenue while its logo disappears from thousands of sites. Both measurements are correct at once.

If you are a vendor, revenue is your number. If you are a buyer wondering whether anyone will answer, footprint is yours.

Does removing chat hurt conversion?

This study cannot tell you. It measures presence, not outcomes. What the data supports is narrower. The widget nobody answered was one that 391 companies of 789 judged not worth keeping between 2021 and 2025. Only 31 were still running it.

By 2026 the configuration was back on 7.0% of the study. On the evidence here the choice worth arguing about is putting a person there against taking it down, and hardly anybody is choosing the first one.

Was 2025 too early to catch the AI agent wave?

In the last version of this article I said that if agent deployment moved the bot-only share materially, the next wave would show it. And that I would publish the figure whichever way it went. Here it is. Bot only went from 12.0% of the study in 2025 to 15.6% in 2026, up 3.6 points. The 95% confidence interval runs 12.1% to 19.0%, which does not rule out no change.

So the answer is that one more year of agent deployment has not clearly moved it, and the wave after this one may. Two things did move. The widget with nobody behind it came back. And the bot handoff collapsed as a share, then delivered a person on 4 of 88 judged sites when we asked it for one.

T
Terry Wilson
Founder, GTM Clarity · CEO, ChatMetrics

Terry Wilson is the founder of GTM Clarity and CEO of ChatMetrics, which has delivered over $5 billion in qualified pipeline and 300,000+ leads for B2B clients across SaaS, services, and industrial sectors. Before founding ChatMetrics, Terry was National Sales & Marketing Manager for a $1B enterprise, leading more than 350 people across Australia. He built GTM Clarity's AI on a corpus of 3M+ real B2B sales conversations that delivered $5B+ pipeline across 200+ companies.

Keep reading