GTMClarity
Method

The website panel, 2021 to 2026, method and dataset

Terry Wilson·October 2026·15 min read

Full figure set: 62 B2B buyer response statistics.

Three waves, and the third one changed what the first two are taken to mean. B2B websites removed chat between 2021 and 2025, and the measure that rose over the same period was a bot offering to bring a person rather than a person being there. The 2026 wave tested that offer by holding a conversation, and it delivered a person on 4 of 88 judged sites.

The articles call it the study. This page calls it the panel, because that is what the design is: the same sites matched across waves rather than three separate samples. That is what lets a figure here say a site removed chat, rather than that fewer sites had it.

This page covers how the panel was drawn, the states each site was classified into, the full results for all three waves, the tests used, and what the design cannot show. The 2021 and 2025 figures were recomputed from source on 29 September 2026 and are exact counts on the matched panel. The 2026 wave is a stratified estimate and its figures carry intervals.

Sample and design

The 2021 wave workbook holds 5,007 rows, of which 5,004 carry a complete single classification, across 4,832 distinct domains. Matching those to the 2025 wave on domain gives the analysis panel of 4,265 sites.

When each wave was collected. The 2021 wave ran in March 2021, as the source workbook describes it. The 2025 wave was collected in April 2025. The 2026 wave ran in September 2026, with coding completed on 17 September 2026.

The months matter for one reason, and it is specific to Drift. Salesloft took the Drift application down on 5 September 2025 and restored it later that month with every playbook and widget returned in the off state, so a Drift site looked at inside that window would have read as having no chat when it had removed nothing. The 2025 wave ran five months earlier, in April, so no Drift figure in this set sits inside that window. The 2026 wave ran a year after it, by which time the application had been accessible for eleven months.

Every panel finding rests on 4,265, and that is the only number a table should carry. The 2021 study published a draw size of 5,006, which matches neither the row count nor the distinct domain count and is not printed as a base.

Each site was above 10,000 monthly visitors, the same traffic floor used across GTM Clarity research, so that findings describe sites with live audiences.

Each site was visited and classified by how a buyer could reach a human. The panel was revisited in 2025 and reclassified on the same scheme. Because the same sites were classified in both waves, the study measures change within sites rather than a difference between two independently drawn samples. A site that removed chat is observed removing chat, not inferred from a shifting aggregate.

Both waves were classified by hand, by a person opening each widget, rather than by a script. Inter-rater drift between the waves is an accepted limitation, as is drift within a single rater over time. Neither can be ruled out from the record this study kept.

The six classification states

Each site was placed in one state.

  1. No chat.
  2. Bot only.
  3. Unstaffed chat, a widget present with nobody behind it.
  4. Staffed chat.
  5. Bot to person.
  6. Bot to unstaffed chat.

Published state names and the coding labels behind them. The 2021 and 2025 workbooks were coded before these names were settled, so two of the six states carry a different label in the source data than on this page. State 5 is coded Bot to Agent and published as bot to person. State 6 is coded Bot to Unmanned Chat and published as bot to unstaffed chat. The data keeps the label it was coded under, because changing a coding value after the fact breaks every derived file that reads it. Map those two before comparing a figure here against the workbook.

Rows carrying more than one state flag are excluded rather than resolved by judgment. Ambiguous classification is the failure mode this design is most exposed to, and excluding those rows keeps every retained row a single unambiguous verdict. The gap between the 5,004 rows with a complete 2021 classification and the 4,265 analyzed is accounted for by domains that could not be matched to a 2025 classification, together with these exclusions.

Human reachable, where it appears in the results, means states 4 and 5 combined, staffed chat and bot to person. It is never published alone. The two states are not equivalent and the difference decides how the whole panel reads.

State 4, staffed chat, means a widget with a person behind it. Its observable meaning is the same in every wave.

State 5, bot to person, means a bot that offered to bring one. The 2021 and 2025 waves classified each site on one visit and recorded the offer. Neither wave tested whether a person arrived, and this page previously carried no operational test for that state. The 2026 wave supplied one, by asking and waiting.

So a combined human-reachable figure spanning 2021 and 2025 mixes a measured state with a promised one. Every result below splits them.

Results

Base for every row: 4,265 sites, in both years.

State202120252026 estimate2026 95% CI2021 to 2025 change
No chat64.0%, 2,731 sites72.3%, 3,083 sites74.1%69.1% to 79.1%+6.7 to +9.8 points, McNemar p below 0.0001
Staffed chat, state 43.3%, 140 sites2.6%, 111 sites2.3%0.4% to 4.1%-1.4 to +0.02 points, p=0.064
Bot to person, state 50.4%, 18 sites11.5%, 490 sites1.1%interval runs below zero+10.1 to +12.0 points, p below 0.0001
Human reachable, states 4 and 53.7%, 158 sites14.1%, 601 sites3.3%0.8% to 5.9%+9.3 to +11.5 points
Unstaffed chat, state 316.7%, 713 sites1.2%, 52 sites6.3%3.2% to 9.5%-16.6 to -14.4 points
Any unstaffed widget, states 3 and 618.5%, 789 sites1.6%, 69 sites7.0%3.8% to 10.2%-18.1 to -15.7 points, p below 0.0001
Bot only13.8%, 587 sites12.0%, 512 sites15.6%12.1% to 19.0%-3.1 to -0.5 points, p=0.009

All percentages are shares of the 4,265 site panel.

The staffed-chat change is not significant. 140 sites to 111 is 0.68 points down with a 95% confidence interval running from 1.38 down to 0.02 up. The panel cannot separate it from no change, and no asset built on this page asserts a fall. What the three waves support is that staffed chat has stayed at about one site in forty throughout.

Chat presence. 753 sites removed chat entirely between 2021 and 2025 and 401 added it, a net movement of minus 352 sites. Discordant pairs 753 and 401 on a base of 4,265.

Where the unstaffed widgets went. Of the 789 sites running an unstaffed widget in 2021:

Outcome in 2025SitesShare of 789
Switched it off39149.6%
Put a person behind it, state 4557.0%
Put a bot in front offering a person, state 517422.1%
Moved to bot only13817.5%
Still unstaffed313.9%

The five destinations sum to 789. Removal beat putting a person behind it by 7.1 to one. 560 of the 789, 71.0%, ended 2025 with no human reachable at all.

An earlier version of this page reported this movement as 390 switched off against 231 staffed, 1.69 to one. That 231 combined the 55 sites that put a person behind the widget with the 174 that put a bot in front offering one. Reporting them together made a promise indistinguishable from a person, and the figure is withdrawn.

The unstaffed widget did not migrate into a single successor state. It resolved in four directions, and the largest of them was removal.

Segment cut. Groups are defined by the 2025 workbook's own B2B and B2C columns, read strictly: a site marked in exactly one of the two is assigned to it, and a site marked in both or neither is excluded. That gives 2,795 B2B sites, 1,368 consumer-facing sites, and 102 excluded.

GroupMeasure20212025
B2B, 2,795 sitesstaffed chat3.11%2.54%
B2B, 2,795 sitesbot to person0.39%12.81%
B2B, 2,795 sitescombined 4 and 53.51%15.35%
Consumer-facing, 1,368 sitesstaffed chat3.80%2.85%
Consumer-facing, 1,368 sitesbot to person0.51%9.58%
Consumer-facing, 1,368 sitescombined 4 and 54.31%12.43%

No significance test is published on this comparison. The direction of the combined measure depends on which of the workbook's sector columns defines the groups, and a comparison that changes sign between two reasonable readings is not a finding. An earlier version of this page published the comparison on two different segment bases with chi-square tests. Those bases cannot be reproduced from the source workbook by any reading of its sector columns, and no derived file holds them, so the figures that sat on them are withdrawn along with the tests.

Consumer sites that appeared in the random draw were retained rather than discarded. That is what makes the comparison a cut within one sample rather than a comparison across two differently constructed samples. It also means the consumer-facing group was not designed as a representative sample of consumer sites, and it should be read as whatever consumer sites fell into a B2B technology draw.

Chat that could reach a person. Across all 1,534 panel sites that had chat of any kind in 2021, 10.3% could reach a person, 158 sites, of which 140 were genuinely staffed.

The 2026 wave

Design. Coded 16 to 18 September 2026. The three large 2025 strata, no chat, bot only and bot to person, were sampled at roughly 110 matched-panel sites each. The three small strata, staffed chat, unstaffed chat and bot to unstaffed chat, were censused. Estimates are weighted to the 2025 stratum sizes on the 4,265 base with a finite population correction. Sampled strata carry a normal-approximation interval at 95%; censused strata carry no sampling error. A share small enough for that interval to run below zero has it reported rather than clipped.

A different instrument, and this is the most important thing on this page. 2021 and 2025 classified a site by opening the widget on one visit and reading what it offered. 2026 held a conversation: a buyer's question, then a fixed three-minute wait, in US business hours. The two instruments answer different questions. The first asks what a site offers. The second asks what a site delivers. Any figure comparing 2026 to an earlier wave states this.

What those words mean, in plain English. A stratum is one of the six states a site could be in at the 2025 wave. The 2026 wave went back to each state separately rather than drawing one sample across the panel, because the states are wildly different sizes and the small ones carry the findings. For the three big states it sampled; for the three small ones it tried every site, which is what censused means. Judged means the 2026 conversation produced a clear answer: some sites were gone and some would not let us in, so the denominator on every 2026 rate is the sites we could actually read, not the sites we tried. This is the canonical definition for the set, and the articles point here.

Within-stratum results.

2025 stateNjudgedstill reached a human in 2026rate95% CI
No chat3,0837922.5%0.7% to 8.8%
Bot only5128222.4%0.7% to 8.5%
Unstaffed chat525147.8%census
Staffed chat1111042221.2%census
Bot to person4908844.5%1.8% to 11.1%
Bot to unstaffed chat1717211.8%census

Of the 104 judged staffed sites, the 82 that no longer reached a human went to unstaffed chat 33, no chat 32, bot only 15, and bot to unstaffed chat 2. So 33 of the 104 judged staffed sites are now running a widget with nobody behind it.

Coding hours. Sites are coded once, in the single 2026 instrument. A site is recoded only where it was tested outside any plausible operating window, which means the small hours. Twelve Australian sites in the staffed block had been coded between 04:33 and 07:59 their own local time and were recoded on 17 September inside Australian business hours; five verdicts changed and the staffed-block result moved from 18.4% to 21.2%. Three European sites in the same block, everest.co.uk, simpson-associates.co.uk and qssolutions.nl, were coded between 20:03 and 22:21 local and are not recoded. Evening is inside the window a widget that is present and offering to answer should cover, and one of the three did reach a human.

Machine detection is a floor. A machine sweep ran before the conversations to find chat. It missed chat on 45 of 155 live sites it had already called empty, 29.0%. Every machine-detected 2026 count of chat presence is therefore a floor. The 2021 and 2025 presence counts were produced by a person opening each widget, not by a sweep, so they are counts rather than floors.

Sites no longer live. 93 sampled domains were not live in 2026 and are excluded from their stratum denominator, which assumes a site that died behaves like the survivors in its stratum. The assumption is stated rather than hidden. 50 of the 93 redirect to a different company, so part of the movement out of chat is consolidation. A further 21 only blocked the sweep and still exist. Within-sample not-live rates were no chat 38 of 120, bot only 28 of 120, bot to person 27 of 120, and 7 of 117 staffed sites unreadable. These are unweighted within-sample rates and no attrition comparison between strata is published on them.

Reading the three waves together

Chat presence fell between 2021 and 2025, so more sites offer no path at all. The measure that rose over the same period was a bot offering to bring a person, not a person being there, and staffed chat did not move. In 2026 the offer delivered a person on 4 of 88 judged sites, the widget with nobody behind it returned to 7.0% of the panel, and four in five of the sites that had been genuinely staffed in 2025 no longer reached a person.

Reported alone, the combined human-reachable rise reads as an industry investing in human contact. It is not what happened, and this page publishes the two states separately for that reason.

Statistics

McNemar's test for paired proportions is used for within-panel change, because the same sites are observed twice and the two observations are not independent. Chi-square tests are used for comparisons between independent groups, such as the B2B and B2C cut. Confidence intervals are reported at 95%.

Relationship to the demo response study

144 domains appear in both the 4,265 site panel and the 1,685 company demo response base, 8.5% of the demo base and 3.4% of the panel, traced by direct domain join. The two studies are close to independent samples of the same market, and a finding from one is not evidence for a finding in the other.

What this panel cannot support

It carries no outcome data of any kind. The panel records what a buyer could see and reach from outside a website. It records nothing about revenue, staffing cost, payroll, conversion or what any company earned from any state. A claim that a staffed site is paying salaries against it, or that a bot-only site is saving money, is not supported here.

It cannot resolve the human reply rate by site type. The demo response study behind the 93.4% figure was powered to measure one overall rate and cannot resolve differences between site types, so no panel segment should be crossed with a never-answered rate.

Vendor attribution is weaker than presence. Presence counts are reliable, because detecting whether a widget exists does not require identifying who built it. Vendor identification fell from 99.5% of sites with chat in 2021, 1,528 of 1,536, to 75.7% in 2025, 896 of 1,183, leaving 287 sites in 2025 where chat was present and no vendor could be attributed. Those denominators are the vendor analysis's own count of sites with chat. The matched panel's count is 1,534 and 1,182, and the vendor numerators have not been recomputed against it, so the detection figures are published on the base they were derived from. Every vendor share figure sits on that rate, and those figures publish as bounds rather than points. The chat vendor share method page carries the detail. None of that caveat applies to the presence and staffing figures on this page.

Observational. The study describes what is visible from outside a website. It records what a buyer could see and reach, not what any company intended.

Silent on what happens next. A site classified staffed chat in 2025 is a site where a person was reachable at the moment of classification. Whether that person answers a given buyer, and how quickly, is the subject of the demo response study and the Buyer Wait Time specification, not this panel.

Point in time. Each site was classified on one visit per wave. A site with intermittent staffing, for example one covering business hours in a single timezone, is recorded as whatever it was at the moment of the visit.

Attrition between waves. Sites that disappeared, were acquired, or could not be classified in 2025 are not in the 4,265. Companies that failed between 2021 and 2025 are absent, which is a mild survivorship effect. In 2026, 93 sampled domains were no longer live and are handled as described above.

Unequal segment sizes. The consumer-facing group is 1,368 sites against 2,795 B2B sites, so consumer-facing figures carry wider intervals and are read as direction rather than as precise measurement.

How to cite this

GTM Clarity, The website panel, 2021 to 2026, method and dataset, 2026. Website data collected March 2021, April 2025 and September 2026. Research directed by Terry Wilson, fieldwork by the GTM Clarity research team.

The findings drawn from this panel are reported in The Buyer Wait Time Report 2026.

T
Terry Wilson
Founder, GTM Clarity · CEO, ChatMetrics

Terry Wilson is the founder of GTM Clarity and CEO of ChatMetrics, which has delivered over $5 billion in qualified pipeline and 300,000+ leads for B2B clients across SaaS, services, and industrial sectors. Before founding ChatMetrics, Terry was National Sales & Marketing Manager for a $1B enterprise, leading more than 350 people across Australia. He built GTM Clarity's AI on a corpus of 3M+ real B2B sales conversations that delivered $5B+ pipeline across 200+ companies.

Keep reading