[Web-bot-auth] Field data on Web Bot Auth from 40,003 origins, two vantages, and a measurement of #93's premise
董文冲 <dwc24@mails.tsinghua.edu.cn> Tue, 11 August 2026 17:38 UTC
Return-Path: <dwc24@mails.tsinghua.edu.cn>
X-Original-To: web-bot-auth@mail2.ietf.org
Delivered-To: web-bot-auth@mail2.ietf.org
Received: from localhost (localhost [127.0.0.1]) by mail2.ietf.org (Postfix) with ESMTP id CFE5A127FA57A for <web-bot-auth@mail2.ietf.org>; Tue, 11 Aug 2026 10:38:15 -0700 (PDT)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=ietf.org; s=ietf1; t=1786469895; bh=CyJGH7tFPUpKpiuC4lwIhYxxKCP6MBzHxGvVl/aXJEg=; h=Date:From:To:Subject; b=xobtUqCWIi5BhK4UvGMiFH4RfwsdhOA3IBy2XTr7ZUVnl8dvFo8BqZRT1OOXybD1o jQUrowUOPOI37QDAxroi9cPtFwmU6zWwrvQnNy1h+ld3zm758SlxQbJimbpKHlO+YQ 0ODJ0k04T5sR3NFQ06F2uIo+k8Dw2nLQnLNcoReI=
X-Virus-Scanned: amavisd-new at ietf.org
X-Spam-Flag: NO
X-Spam-Score: -1.697
X-Spam-Level:
X-Spam-Status: No, score=-1.697 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_INVALID=0.1, DKIM_SIGNED=0.1, RCVD_IN_MSPIKE_H2=0.001, RCVD_IN_VALIDITY_RPBL_BLOCKED=0.001, RCVD_IN_VALIDITY_SAFE_BLOCKED=0.001, SPF_HELO_NONE=0.001, SPF_PASS=-0.001] autolearn=no autolearn_force=no
Authentication-Results: mail2.ietf.org (amavisd-new); dkim=fail (1024-bit key) reason="fail (body has been altered)" header.d=mails.tsinghua.edu.cn
Received: from mail2.ietf.org ([166.84.6.31]) by localhost (mail2.ietf.org [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id nZtzLp5P0rjG for <web-bot-auth@mail2.ietf.org>; Tue, 11 Aug 2026 10:38:14 -0700 (PDT)
Received: from azure-sdnproxy.icoremail.net (azure-sdnproxy.icoremail.net [13.75.44.102]) by mail2.ietf.org (Postfix) with ESMTP id 3098A127FA574 for <web-bot-auth@ietf.org>; Tue, 11 Aug 2026 10:38:13 -0700 (PDT)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=mails.tsinghua.edu.cn; s=dkim; h=Received:Date:From:To:Subject: Content-Transfer-Encoding:Content-Type:MIME-Version:Message-ID; bh=wCjp3AfLULGSuZd2818KBPxQKvaW/OLEA0cYZA54muY=; b=esjhO4nRgxSdZ c/vlM/ujijAxBzpKdobrVtcuX4F24yIX3uJWS0DpimM3LYsuZw7mjHAyRcb+NJoa KU69696G0xJxB9ewfl1Dj+Izh5/aY6o8nK73EPiixPcRrmWM7XvZd9aO00ba0h0K GVk+RxAWABRk3DMYgN7jJv2JvrJ5Ok=
Received: from dwc24$mails.tsinghua.edu.cn ( [121.35.182.72] ) by ajax-webmail-web1 (Coremail) ; Wed, 12 Aug 2026 01:38:10 +0800 (GMT+08:00)
X-Originating-IP: [121.35.182.72]
Date: Wed, 12 Aug 2026 01:38:10 +0800
X-CM-HeaderCharset: UTF-8
From: 董文冲 <dwc24@mails.tsinghua.edu.cn>
To: web-bot-auth@ietf.org
X-Priority: 3
X-Mailer: Coremail Webmail Server Version 2024.2-cmXT5 build 20250909(015d6f0a) Copyright (c) 2002-2026 www.mailtech.cn mispb-4df55a87-4b50-4a66-85a0-70f79cb6c8b5-tsinghua.edu.cn
Content-Transfer-Encoding: base64
Content-Type: text/plain; charset="UTF-8"
MIME-Version: 1.0
Message-ID: <1f207562.193b4.19ff1e73b2b.Coremail.dwc24@mails.tsinghua.edu.cn>
X-Coremail-Locale: zh_CN
X-CM-TRANSID: yAQGZQBnvMsCXntqEhGsAQ--.39433W
X-CM-SenderInfo: hgzfjko6pdxz3vow2x5qjk3toohg3hdfq/1tbiAgYQEWp7IblUVQA CsP
X-Coremail-Antispam: 1Ur529EdanIXcx71UUUUU7IcSsGvfJ3iIAIbVAYjsxI4VWkKw CS07vEb4IE77IF4wCS07vE1I0E4x80FVAKz4kxMIAIbVAFxVCaYxvI4VCIwcAKzIAtYxBI daVFxhVjvjDU=
Message-ID-Hash: HRRUPDLQYMHW5Z2DEXT5J3QT4BCL4OXQ
X-Message-ID-Hash: HRRUPDLQYMHW5Z2DEXT5J3QT4BCL4OXQ
X-MailFrom: dwc24@mails.tsinghua.edu.cn
X-Mailman-Rule-Misses: dmarc-mitigation; no-senders; approved; emergency; loop; banned-address; member-moderation; header-match-web-bot-auth.ietf.org-0; nonmember-moderation; administrivia; implicit-dest; max-recipients; max-size; news-moderation; no-subject; digests; suspicious-header
X-Mailman-Version: 3.3.9rc6
Precedence: list
Subject: [Web-bot-auth] Field data on Web Bot Auth from 40,003 origins, two vantages, and a measurement of #93's premise
List-Id: Authentication of non-human users to human-oriented Web sites <web-bot-auth.ietf.org>
Archived-At: <https://mailarchive.ietf.org/arch/msg/web-bot-auth/10wyDxpZs4MyWgkMkwb3kVmefSQ>
List-Archive: <https://mailarchive.ietf.org/arch/browse/web-bot-auth>
List-Help: <mailto:web-bot-auth-request@ietf.org?subject=help>
List-Owner: <mailto:web-bot-auth-owner@ietf.org>
List-Post: <mailto:web-bot-auth@ietf.org>
List-Subscribe: <mailto:web-bot-auth-join@ietf.org>
List-Unsubscribe: <mailto:web-bot-auth-leave@ietf.org>
Hi all,
Thibault suggested I bring this to the list.
I am a PhD student at Tsinghua. For an academic measurement study I implemented the signing profile by hand, published a directory at agent.hiem.al, and ran signed and unsigned arms against a 40,003-origin frame from Tranco. This is written against draft-meunier-webbotauth-httpsig-protocol-01 of 6 August, and I have read the review opened on 7 August, so I know some of it is live. Issue numbers below are on the draft's repository: #93 is the FCrDNS fallback, #111 is "can't validate without a key", #104 and #106 are the closed enrolment issues.
What this is not. It is not a conformance report, and I do not think "the mechanism is not deployed yet" is worth anyone's time: the document is a draft, adoption takes years, and absence of deployment at this stage is what you would expect. Three things I can offer instead, none of which depends on treating the current text as fixed. What it costs a signer to follow the draft today. What one specified mode does in the field. And a measurement of the premise under #93, which is about FCrDNS and not about this draft at all.
Method, briefly, because three details decide how much of this you can believe. The frame is 40,003 origins from Tranco, stratified by rank. The client is a plain HTTP client and not a browser: no JavaScript, one TLS fingerprint throughout, and the arms are byte-identical except for the headers under test, so an arm cannot differ from another by anything but its declaration. The addresses are datacenter ones in Los Angeles, London and Tokyo, which I expect is itself part of any bot score I am measuring. robots.txt is fetched and obeyed per identity. "Admitted" means the origin returned at least 100 words; "refused" means fewer than 20, or a 4xx or 5xx; between those I record undecided, and undecided enters no comparison.
Signing, un-enrolled, costs access. The arms are paired, so the same origins carry all of them and the discordant cells are the evidence; marginal counts are context, not the test. Against an otherwise identical unsigned request, adding Signature, Signature-Input and Signature-Agent lost admission at 657 origins and gained it at 34, of 26,548 where both decided. An earlier pass over the same frame, on the instrument we used before 5 August, gave 408 against 38, so the sign is not one run's fluke. Nobody is checking the signature: three unfamiliar headers are moving a bot-management score.
Since the private note, I have run the same arms from a second vantage, because a measurement from one address invites the reading that it measured the address. This is the part I would most want a list to push on. On 11 August I re-ran the arm set from a different host, in a different country, on a 6,000-origin subset of the same frame, holding everything else identical -- same arms, same repeats, same ceilings, same delays, same library versions. Compared on those same 6,000 origins:
| | first vantage, 6 Aug | second vantage, 11 Aug |
|---|---:|---:|
| lost admission when signed | 96 | 117 |
| gained | 2 | 8 |
| of origins deciding both | 4,021 | 4,035 |
| discordance | 2.44% [2.00, 2.96] | 3.10% [2.61, 3.68] |
The intervals overlap and both are far past any conventional threshold. What makes this more than a repetition is the contrast between the two rows of the experiment. Between those two runs the individual arms disagree with themselves about 5.0-5.5% of origins -- vantage plus five days of ordinary churn -- while the signed-versus-unsigned difference stays put. On one address a day apart that same self-disagreement is 0.5%, so I know what the floor looks like. The level of admission is a property of who is asking and when; the cost of signing is not. That is what I would expect if the effect is real, and it is the check I would have demanded of someone else's single-vantage number.
Everything else replicates on the second vantage too. Valid and corrupt signatures remain indistinguishable: 6 origins against 2, p = 0.29. The cell an origin that actually verifies would have to land in -- valid admitted while corrupt is refused -- holds 2 origins, against 6 in the incoherent mirror of it, so it is negative on net and indistinguishable from noise. And signing costs 10.4 times this run's own in-frame floor.
I want to be careful about what that does and does not say, because I think it says less than it first looks like. It is a measurement of signing while un-enrolled, which is not a measurement of what Web Bot Auth does. An enrolled crawler is presumably recognised and I cannot observe one: none of the crawler identities I impersonate and measure signs anything at all and I cannot acquire a verified key, so there is no enrolled signer anywhere in my data. Nor can I separate the two mechanisms that would produce this result, and they point in opposite directions. If unfamiliar headers are simply scoring badly, the effect should fade as the headers become common and it says little about the design. If instead a signature that fails an allowlist lookup is treated worse than no signature at all, that is a property of the gate and it will not fade. My unsigned arm is the same unknown crawler, so what follows is attributable to the headers and not to being unknown, but which of those two readings applies is not something this experiment can decide.
On #111, "can't validate without a key", I happen to have run exactly that mode. The spec's position is that Signature-Agent is RECOMMENDED but not required and that verification still works where the verifier already holds the key. My no-Signature-Agent arm is that mode: a valid signature over ("@authority"), keyid thumbprint as the only identifier. On 6 August, over the same 40,003 origins:
| arm | admitted |
|---|---:|
| unsigned | 18,501 |
| signed | 17,893 |
| signed, corrupt | 17,904 |
| signed, unlisted key | 17,882 |
| signed, no Signature-Agent | 17,823 |
It is not neutral and it is not merely equal. It is the worst of the five, and against signed-with-header the difference is significant: 69 origins admit the header form where no-agent is refused, against 29 the other way. I offer that as data for the thread and not as a view on how the issue should close.
Two caveats, and the first is the one I would raise if I were reading this. The four signed arms sit between 17,823 and 17,904 admitted, a spread of 81 on a 40,003-origin frame, which is 0.20%. Paired, the discordant cells among them are small and all of one order: valid against corrupt is 30 to 31 on 5 August and 38 to 18 on 6 August; corrupt against unlisted-key is 29 to 35 and then 36 to 14; valid against unlisted-key is 25 to 31 and then 26 to 23. Two of those move between runs in a way I cannot separate from noise, and one of them, 38 to 18, reaches p = 0.011 in the direction of the corrupt signature doing better. With six pairwise comparisons in that run the threshold is 0.0083, so it does not survive, and I am not going to build anything on it.
I am also not going to resolve it, and the reason is ethical rather than technical. Settling a difference this small would mean repeating the same six-arm pass over the same 40,003 origins several times in a short window, and that is a real imposition on real sites for a question that is secondary to everything else here. So I am reporting the numbers and stopping there.
The second caveat is my own noise floor, and here the direction matters. I measure noise as corrupt against unlisted-key, both invalid and so one stimulus to an origin that does not verify: 0.19% on 6 August, 0.24% on 5 August. But on 6 August those two separate at 36 to 14, p = 0.003, which does survive the correction and which they should not do if both are simply invalid. One candidate is the keyid string itself. My unlisted-key arm sends keyid="unpublished", eleven characters, where a real keyid is 43 of base64url. If that is what is being scored then my floor is inflated and every multiplier I quote against it is an understatement. I would rather tell you that than have you find it.
The part I would most like views on is #93. I want to be careful about its state: it is open, it has no resolution, and what it proposes is that implementers MUST fall back to existing verification techniques. I am not describing that as settled. What I can add is a measurement of its premise, because a fallback is only worth mandating if the thing being fallen back to is performed. From two hosts whose PTR fails Google's published procedure, of 1,777 paired origins that served a browser, 1,632 admitted the Googlebot claim at one host and 1,632 at the other, 91.84% at each; 1,631 of them, 91.78%, admitted it at both, and the hosts disagreed about two origins.
This is not a criticism of the procedure or of Google. Google's own procedure works: 12 of 12 addresses I sampled forward-confirm and the PTR names Google. The difficulty is that it is Google's, and for most of the crawlers a fallback would have to cover there is nothing to fall back to. Of the ten identities I could sample, two have a PTR that names the operator, both of them Google's. Three have no PTR at all: GPTBot, OAI-SearchBot and ChatGPT-User, where a reverse lookup returns nothing to check. The remaining four forward-confirm and name the hosting provider instead of the operator, which is the case I would most want a sentence about: an ec2 hostname of the compute-1.amazonaws.com form round-trips perfectly, is reproducible by anyone who rents an instance in that range, and passes an FCrDNS check while identifying nobody. So a verifier that falls back to FCrDNS gets a usable answer for two identities, a false pass for four, and nothing at all for three.
Put beside the percentage, that is what I would offer to #93 rather than a position on it. Origins largely do not run the check, and for most crawlers there is not much of a check to run. If the MUST lands, a sentence saying that a forward-confirmed PTR may name the hosting provider rather than the operator would stop an implementer reading a pass as an identification. Requiring attribution to a domain the operator controls is the fix email settled on.
One observation that supports the above rather than standing on its own. Across 294,898 signed requests no origin fetched our published key, and no response in 3,447,202 carried Accept-Signature. I do not offer that as a finding. Nobody fetching an unknown signer's directory is rational behaviour, not a gap, and a two-year-old draft not being widely deployed is what anyone should expect. What it does show is which path is actually in use: a production verifier answered "401 unknown public key or unknown verified bot ID for keyid", parsing and rejecting on absence without fetching. That is a key allowlist, and an allowlist is an enrolment gate under another name. A positive control says these origins were deciding on identity throughout: one AI crawler identity against a browser identity is 2,592 discordant against 262.
And the enrolment path, which I raise knowing #104 and #106 are closed. I am not asking for them to be reopened; what I have is a measurement of the deployed consequence rather than a comment on the text, and it was not available while they were open. Reputational continuity is the right thing to build on, but nothing takes a signer from correct to known except a per-verifier application, which is the scalability problem the draft says IP allowlisting has, reproduced with keys. One narrow version of it selects on the wrong axis: an enrolment gate running through an organisational account is easy for a company and awkward for a PhD student, who cannot enter their university into a programme agreement. A line about who may enrol, and at what granularity, would help the people least able to comply. For context, we applied to one such programme on 31 July and have had no decision; the endpoint still answers 401. That is not a complaint and I am not chasing it, because a decision that arrives on its own makes the before-and-after honest.
Three questions for the list. Does the un-enrolled signing cost belong anywhere in the draft as a deployment consideration, or is it out of scope for a protocol document? For #93, would a sentence distinguishing a forward-confirmed hosting-provider name from an operator identification be welcome, and where? And is there anything about the enrolment path itself that the drafts should say and currently do not?
Happy to send the full numbers, the exact signature base, the directory, or the read-out code. The measurement is not finished and nothing above is load-bearing on a result I would mind correcting in public.
Best,
Wenchong Dong