[Web-bot-auth] Re: Follow-up from IETF 126: Comments on draft-illyes-webbotauth-cbcp

Joshua Ashcroft <josh.ashcroft@gmail.com> Mon, 24 August 2026 06:17 UTC

Return-Path: <josh.ashcroft@gmail.com>
X-Original-To: web-bot-auth@mail2.ietf.org
Delivered-To: web-bot-auth@mail2.ietf.org
Received: from localhost (localhost [127.0.0.1]) by mail2.ietf.org (Postfix) with ESMTP id 76B2012E2FCA7 for <web-bot-auth@mail2.ietf.org>; Sun, 23 Aug 2026 23:17:00 -0700 (PDT)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=ietf.org; s=ietf1; t=1787552220; bh=Q8ZduqzHH5e51fBZsod0tKYNnKGUZvpDIYvQ5vZRJ58=; h=References:In-Reply-To:From:Date:Subject:To:Cc; b=VbrJLN2AHqAFVHlIcU9HLtKrGRGVFyuWeNMGpo+nDH8yS1iXqxgIRYwS8SuN3KZxT saX/5ZJer/ElUTbBdrDJLYgic0SqkQQTwHDJN535PJwTN8KoDbiZHE0HRLg077VD96 z9EvG7a8CdVaCqSmP3Q1pbzlCcj637QIdUZ6tn8s=
X-Virus-Scanned: amavisd-new at ietf.org
X-Spam-Flag: NO
X-Spam-Score: -2.098
X-Spam-Level:
X-Spam-Status: No, score=-2.098 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, FREEMAIL_FROM=0.001, HTML_MESSAGE=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_HELO_NONE=0.001, SPF_PASS=-0.001] autolearn=ham autolearn_force=no
Authentication-Results: mail2.ietf.org (amavisd-new); dkim=pass (2048-bit key) header.d=gmail.com
Received: from mail2.ietf.org ([166.84.6.31]) by localhost (mail2.ietf.org [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id GX4X-emPe8rw for <web-bot-auth@mail2.ietf.org>; Sun, 23 Aug 2026 23:16:59 -0700 (PDT)
Received: from mail-yx1-xb130.google.com (mail-yx1-xb130.google.com [IPv6:2607:f8b0:4864:20::b130]) (using TLSv1.3 with cipher TLS_AES_128_GCM_SHA256 (128/128 bits) key-exchange X25519 server-signature ECDSA (P-256) server-digest SHA256) (No client certificate requested) by mail2.ietf.org (Postfix) with ESMTPS id 42DB912E2FCA0 for <web-bot-auth@ietf.org>; Sun, 23 Aug 2026 23:16:59 -0700 (PDT)
Received: by mail-yx1-xb130.google.com with SMTP id 956f58d0204a3-66d07ddf077so781441d50.2 for <web-bot-auth@ietf.org>; Sun, 23 Aug 2026 23:16:59 -0700 (PDT)
ARC-Seal: i=1; a=rsa-sha256; t=1787552212; cv=none; d=google.com; s=arc-20260327; b=g/tNwSxwEAo/kKeAvhRAi028HIVk38kYUPV6RNFHMBKvn+j6viaiPb7E8DvjCwAVNi peGSIr3C8lSIbHP7Dt1XcpjfWbMXux8ZflRu8DbiUHImmF1vwsUr4+PR0Z8ZXmcXbgRx IoCuRKWt3FLsYhgNjrII3BPH7CtWTTt7qhS4flk4c3ioyrsD6h/Tcj/H/8jypPp2CB01 BZoZ2HwUOjpcxDjRP7+uomGiIBo4W2nEYEUPW9fXakm69kd4LFkjRItnmv/Z4vR+sYSQ 6LxYlLBFgtfrZOZU76sUu+gFRvBG3uDKqaUZvB444OBDO5qlz5uFi2BOfPHqOKkbZVjL R6jA==
ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=arc-20260327; h=cc:to:subject:message-id:date:from:in-reply-to:references :mime-version:dkim-signature; bh=ZW0VqOthYxihJrKzl59PShL2AYmfdqAql7vka3Cdp38=; fh=I7qPd7QB2KEejJsXh83Bh5LtojZyKe74ZeUjTu3Ibl0=; b=YjEkr1cEXwx/TV3wZoHP2YqODZCBXr3cK/FqT676EJyHWNhSwALEMnhk7GBu3lNc9W kdPOcg2jQ40jVN4TQAKBH0S6FbhdxYIATAOmsBbANciq5aJfJdzojG01jczitmJO7xFb qOQEFMO/Fvy1XocEWVTJCtQlynBj+tS7F7eZbrsphnOq5F6SXlAEl3VOIWCv+10VBXmw YiZDhsklsh68mzHigAUV9Dv6f5kmxQm27ee91dP8wZeFZH4JSF71W65bL8Bt6TyCUVj9 dRhZFKEW5/Fb7PmKo+/7wje3YuhpBws8301kzTP97Av+yf5OsSgu52Ef95ysNbuqAK9a Swcg==; darn=ietf.org
ARC-Authentication-Results: i=1; mx.google.com; arc=none
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787552212; x=1788157012; darn=ietf.org; h=content-type:cc:to:subject:message-id:date:from:in-reply-to :references:mime-version:from:to:cc:subject:date:message-id:reply-to :content-type; bh=ZW0VqOthYxihJrKzl59PShL2AYmfdqAql7vka3Cdp38=; b=kQUJynaUZyXFS7wfHYdPrA5Xj5AnEnhapEddo2xzrUJOF9ZgNq/4/rbUNJL+/QiBLu ZqTCA+q4yIQ+W041mzjiaQLhKPxb0dFXdSmhlTAVYLTAUlcvAaS2DwY8HOv4HyQcSQXH DYvsMXv8iLCKqueFgoVdYX6vjbeBD7VCPz34khEllbdCfmv03R+iKhFdUc3NtrflhNKa 8kFZl8S8jpD2WVkhUc8437UkFUPoSwRhSrhR0xc34QWNgnN0Pu8BGyfKCs/3tWzRLjQm tggbk9poA0iRqQCD5sP21uziM4tMKL3Xn2uL9rA5k+0V4hxwn9HeG6QsIZIvcwKwxejq PfxQ==
X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787552212; x=1788157012; h=content-type:cc:to:subject:message-id:date:from:in-reply-to :references:mime-version:x-gm-gg:x-gm-message-state:from:to:cc :subject:date:message-id:reply-to:content-type; bh=ZW0VqOthYxihJrKzl59PShL2AYmfdqAql7vka3Cdp38=; b=D3LVPNsz+GDtlT4bzpc0CHvY4+4krZL4yczAGVGbTVD7eypgkXPndYLzrbf7UmNpDN P0zHch9gEgk9J0HyyRTtF4irmrgVhe7be02bv53x0htB5etVw472PYr6WLpSYVN2NW5M 5sHyFAYEeeQt24DAdiSM+H/A6moPFexyTXxPMmFSplRbfD4niICybItT6Ae9j/ky8b8W feGF27xEjoC3VDrkeEGpnglEPNDYeFgowp1W/v60uHCR+NVBlQT3mlQ+8k4oosujSsla C6pcGlzebfeQ015bhnm0vRWXbSOZqHe3OTrh8ykWOVDvnz6RyaKc6M5BSnf2FnLGC9Lx WFcg==
X-Gm-Message-State: AFuF++k1LBbOU6dJlWW0vmWGvv9lMf+C6te0IOE2/ASkwN3FgFxcrCud wVsxbS/8K5CqUN9aUCrPvN9k5Ak4gxR0n2NiCQrS/1Kc8jyP4UGIIFg5JbII9Na37k52pnIyT8V h1WsD5nkwP9DY/gamib6Yp+RGW+tp05o=
X-Gm-Gg: AR+sD10sohD4olGciod/P8kuZ6jRh2CE3Eh7zdhJ8QZEZFFQVNM0UW6p5PkzMLalXY7 pny1Lr6456CQOp3ofhfj1ryCKV4Vqo+8j35S6dcD/lrjaU7LZt55XC8S+by1vcgHVTZhuGUtgaf InDt+kBI9aIMZ7W+XanXTn8Ay5+MbgxrtMtpEQz2Fn+flzcrhxbVvZvq9tlpK57jL7/49qMdX9p hDqJMPf8288XB0HgthhPHSRt6eA5NTGFlkO/HIUjyf+b8Ie/m4lm80ulRwfJ346vpjfJK6RfHZR +tOPwOjgBSNINRVLL8vTkeFey/YXyf/OGmdLcKzRhOCpnMeINbcTfn+83HkxU96ectmvILLzBw= =
X-Received: by 2002:a05:690e:450b:20b0:66c:e3db:e14e with SMTP id 956f58d0204a3-66cf213822dmr3255787d50.8.1787552212406; Sun, 23 Aug 2026 23:16:52 -0700 (PDT)
MIME-Version: 1.0
References: <CADTQi=dSGTGJtczLTZ1ua77C8h4pvGhrVkz3MZV7SO280MJrKQ@mail.gmail.com> <19B52E65-D016-4606-A329-70D6FF10E955@vaibhavbajpai.com> <CAKaUSfUFR82mCjy7qWzgibnuatQ7YdJ_yStur3b2zMR0KbMWkQ@mail.gmail.com> <6d65e967-b8a0-46be-9d02-00b66adb306a@Spark> <CAKaUSfUH=yZiHgP0=YjCnEpFKyVC3meZV9MYH+JfjYqQcF1sfA@mail.gmail.com>
In-Reply-To: <CAKaUSfUH=yZiHgP0=YjCnEpFKyVC3meZV9MYH+JfjYqQcF1sfA@mail.gmail.com>
From: Joshua Ashcroft <josh.ashcroft@gmail.com>
Date: Sun, 23 Aug 2026 23:16:39 -0700
X-Gm-Features: AcwNN1WjxTXXyCewwutlCTrEjRlKgr3HBif6wBSiLPTSEPikI4psCm1hJ86zPOQ
Message-ID: <CAJ9=2ZsHLqAAzUMLZQUNzYh2FQH2LOW49SiWJsU8C+YUOa2b7w@mail.gmail.com>
To: shivdeepsachdeva@gmail.com, nick@lifelightlabs.com, contact@vaibhavbajpai.com
Content-Type: multipart/alternative; boundary="000000000000a3546f0659c4eed4"
Message-ID-Hash: L7NMCEKBHPR7TDLMX7XWE6FUDOBRWWJ7
X-Message-ID-Hash: L7NMCEKBHPR7TDLMX7XWE6FUDOBRWWJ7
X-MailFrom: josh.ashcroft@gmail.com
X-Mailman-Rule-Misses: dmarc-mitigation; no-senders; approved; emergency; loop; banned-address; member-moderation; header-match-web-bot-auth.ietf.org-0; header-match-web-bot-auth.ietf.org-1; nonmember-moderation; administrivia; implicit-dest; max-recipients; max-size; news-moderation; no-subject; digests; suspicious-header
CC: web-bot-auth@ietf.org, sauron@google.com, ietf@kuehlewind.net
X-Mailman-Version: 3.3.9rc6
Precedence: list
Subject: [Web-bot-auth] Re: Follow-up from IETF 126: Comments on draft-illyes-webbotauth-cbcp
List-Id: Authentication of non-human users to human-oriented Web sites <web-bot-auth.ietf.org>
Archived-At: <https://mailarchive.ietf.org/arch/msg/web-bot-auth/A59l1HBa-4HwoG-97TGwGjDQbFA>
List-Archive: <https://mailarchive.ietf.org/arch/browse/web-bot-auth>
List-Help: <mailto:web-bot-auth-request@ietf.org?subject=help>
List-Owner: <mailto:web-bot-auth-owner@ietf.org>
List-Post: <mailto:web-bot-auth@ietf.org>
List-Subscribe: <mailto:web-bot-auth-join@ietf.org>
List-Unsubscribe: <mailto:web-bot-auth-leave@ietf.org>

Shivdeep, Nick, Vaibhav,

Second implementer datapoint on the scope question first: we implement
B.1 on directory-type responses only, same reading as Nick's verifier,
for the same reason - it is what the text licenses and what the vectors
cover. So a production verifier and now two production directories have
independently landed on the narrow reading, which makes the ambiguity
worth a sentence whichever way it resolves.

One consequence I have not seen raised: the scope decision has an
enrollment multiplier. Cloudflare's enrollment requires a signature of
your directory, the B.1 artifact, as a condition of registration. If
B.1 is directory-only, then operators pushed to jwks_uri or cimd are
not just losing domain binding and rotation per 4.3 - they cannot
produce the artifact the flagship enrollment gate requires. The BCP
would be pointing the operators least able to comply at a mechanism
that locks them out of the program that makes signing worthwhile in the
first place. Whichever scope is intended, the sentence belongs where
those operators will read it.

Shivdeep, your four-under-one deployment makes concrete what I was
about to describe hypothetically, so let me answer the
origin-per-operator cost from the other host seat. We run per-identity
origins, and the marginal cost has been a wildcard certificate and a
routing rule: a subdomain allocated once at onboarding, directory type
at its well-known path, rotation and B.1 intact. The same move covers
Vaibhav's per-purpose point, since a distinct purpose is just another
subdomain. -02's redirect prohibition pushed us into the adjacent shape
last week - each authority serving its own 200 with its own proof - and
that migration was a routing rule and a second proof, not new
infrastructure (details in my implementer's report from Aug 19). So
from where I sit the origin requirement is real but cheap.

The honest limit is a different one, and it may deserve a sentence
wherever the scope sentence lands: a minted origin is an origin under
the host's domain, so the 4.5 domain binding names the host, not the
operator. The operator gets directory-type mechanics - rotation, B.1,
the enrollment artifact - but the name those bind to is mine. Whether
that is acceptable depends on what domain binding is for, and the
answer is probably policy for the verifier rather than a rule here,
but the asymmetry should be visible rather than discovered.

And yes to your offer: if Thibault lands on directory-only, your
sentence in the place a jwks_uri or cimd operator would actually find
it is the right fix, and a PR beats leaving it with the editors. I
second Nick's ask for the scope confirmation before the BCP points
operators at the mechanism.

On the merits of the scope itself: as Nick observed, nothing in the
B.1 construction consumes a directory-format field - it applies
mechanically to any HTTPS response that serves key material - and a
proof mechanism gated on controlling an origin selects for who can
pay, not who holds the key, which is the same scaling problem this
protocol exists to retire.

Joshua Ashcroft

On Fri, Aug 21, 2026 12:17 AM, Shivdeep Singh <shivdeepsachdeva@gmail.com>
wrote:

> Nick, thank you, and the tag is the sharper version of this. I had the
> scope sentence and the consumption path in 5.5.3. That a conformant proof
> must carry tag=http-message-signatures-directory is what makes it concrete,
> since that value has no story on a jwks_uri response.
>
> Agreed that the ambiguity is the problem rather than either reading.
>
> This is not hypothetical at our end. We serve a hosted directory carrying
> four operators' keys under one identifier. Per-tenant authorities would
> resolve that, but they require an origin per operator, so the route for
> operators without one is per-tenant jwks_uri, which lands squarely in the
> gap.
>
> So if Thibault lands on directory-only, I am happy to draft the sentence
> for wherever a jwks_uri or cimd operator would actually find it: that key
> material published under those types cannot carry an Appendix B proof, and
> that a verifier receiving it redistributed falls back to the thumbprint
> identifier in 4.3, which does not survive rotation. Happy to open it as a
> PR, or leave it with the editors.
>
> Shivdeep
>
> On Fri, Aug 21, 2026 at 8:49 AM Nick Mathews <nick@lifelightlabs.com>
> wrote:
>
>> Shivdeep, Vaibhav, Gary,
>>
>> Speaking to the Appendix B question as the contributor of the test
>> vectors that validate it (E.2.3), not for the spec text itself.
>>
>> I think both of your readings have support in the text as written, which
>> is the problem. The framing sentence describes the concern: what a verifier
>> checks when it wants the domain a key is published under, rather than the
>> URL on its own. Nothing in the B.1 construction consumes a directory-format
>> field; it is a per-key signature over @authority and content-digest with
>> created and expires, and that applies mechanically to any HTTPS response
>> that serves key material, including a hosted jwks_uri. But the normative
>> hooks are directory-shaped: the required tag value is literally
>> http-message-signatures-directory, and the consumption path runs through
>> 5.5.3. So an implementer today cannot claim conformant proofs on a jwks_uri
>> response even though the mechanism would work there.
>>
>> Our verifier implements B.1 on directory responses only, because that is
>> what the text licenses and what the vectors cover. If the intent is
>> any-type, the appendix should say so and give the tag a story for
>> non-directory responses. If the intent is directory-only, then the
>> consequence you describe, jwks_uri and cimd operators falling back to
>> thumbprint identity with no rotation per 4.3, deserves a sentence where
>> those operators will find it. Either sentence is cheap and the ambiguity is
>> not. I would like to see Thibault confirm the intended scope before the BCP
>> points operators at the mechanism.
>>
>> On the point 3 convergence, one merchant-side datapoint: distinguishable
>> per-purpose identities are not just registry hygiene. Our policy layer keys
>> rules on the resolved identifier, so a crawler that separates search from
>> training from assistant retrieval gets separate policy rows a merchant can
>> treat differently. A single generic identity collapses that to one row, and
>> the measurement in the paper matches what we see from merchants: they want
>> to say yes to one purpose and no to another, and the identifier is the only
>> handle they have.
>>
>> Nick
>> On Aug 20, 2026 at 6:26 AM -0400, Shivdeep Singh <
>> shivdeepsachdeva@gmail.com>, wrote:
>>
>> Vaibhav, Gary, Nick,
>>
>> Point 3 and Nick's argument converge on the same mechanism, and I have a
>> scope question about it.
>>
>> Under draft-meunier-webbotauth-httpsig-protocol-02 an identity is the
>> resolved Signature-Agent URL, so distinguishable per-purpose crawlers mean
>> distinct URLs. An operator without a stable origin cannot use the directory
>> type, since 5.5 requires that member value to be an origin, which leaves
>> jwks_uri or cimd.
>>
>> Appendix B opens with "It applies to the directory type in Section 5.5."
>> I read that two ways. Either the possession proof in B.1 is available only
>> to the directory type, in which case jwks_uri and cimd material can never
>> satisfy 5.5.3 and falls back to thumbprint identity, which 4.3 says has no
>> rotation. Or the scope line is about domain binding specifically, the
>> concern of that appendix, and B.1's mechanism is available to any type, in
>> which case saying so would help implementers.
>>
>> I do not think this changes either recommendation. It affects who can act
>> on them, so it seemed worth asking before the BCP points operators at the
>> mechanism.
>>
>> Shivdeep Singh
>>
>> On Thu, Aug 20, 2026 at 2:53 PM Vaibhav Bajpai <contact@vaibhavbajpai.com>
>> wrote:
>>
>>> Hi everyone,
>>>
>>> This paper just got published, and we believe would be
>>> valuable input to the BCP draft:
>>>
>>> >From robots.txt to ai.txt: Mapping the Evolution of Web Permissions in
>>> the Age of AI
>>> https://dl.acm.org/doi/pdf/10.1145/3831956.3831960
>>>
>>> The measurements suggest that being identifiable and respecting
>>> robots.txt is
>>> necessary, but no longer sufficient. The harder emerging problem is
>>> ensuring
>>> that authenticated crawlers interpret a fragmented set of publisher
>>> signals
>>> consistently, predictably, and according to the intended type of AI use.
>>>
>>> Our measurements show that today’s system combines heavy dependence on
>>> robots.txt,
>>> fragmented AI-specific policies, inconsistent naming, and only partially
>>> coherent configurations.
>>>
>>>
>>> To this end, here is some input for the BCP:
>>>
>>> 1.) Define how crawlers should handle conflicting signals and precedence:
>>> Explicitly establish that discovery/descriptor files do not grant access
>>> and
>>> recommend conservative handling of contradictory permission signals.
>>>
>>> 2.) Separate crawler identity by purpose. Search, training, and
>>> user-triggered/assistant retrieval should be distinguishable where they
>>> represent materially different uses. Measurements show that that
>>> publishers
>>> are expressing use-specific preferences: e.g., search=yes, ai-train=no.
>>>
>>> 3.) Strengthen identity requirements and encourage a canonical crawler
>>> registry: Documentation alone isn’t solving crawler identity, as is
>>> apparent
>>> from the very large number of unrecognized User-Agent strings in our
>>> measurements. Plus, when an operator runs crawlers for materially
>>> different
>>> purposes, those crawlers should have distinguishable identities, e.g.:
>>>
>>>   ExampleSearchBot
>>>   ExampleTrainingBot
>>>   ExampleAssistantFetcher
>>>
>>>   rather than one generic ExampleBot.
>>>
>>> 4.) Explicitly prohibit identity evasion/policy shopping: A crawler
>>> shouldn’t
>>> spoof identities, rotate identifiers to evade rules, or choose whichever
>>> applicable policy gives it the most permissive result.
>>>
>>> 5.) Encourage policy revalidation rather than indefinitely caching
>>> permission
>>> decisions: Crawlers should periodically revalidate cached crawler-policy
>>> resources and must not assume that a previously observed permission
>>> remains
>>> valid indefinitely.
>>>
>>> 6.) Require transparent interpretation and validation. Crawler operators
>>> should
>>> document supported policy mechanisms, conflict behavior and caching, and
>>> ideally provide a URL testing tool showing publishers exactly how their
>>> crawler interprets the site’s policies, e.g., “This is how our crawler
>>> interprets your site.”
>>>
>>> Kind regards, Vaibhav
>>>
>>> > On 27. Jul 2026, at 14:42, Sauron <sauron=40google.com@dmarc.ietf.org>
>>> wrote:
>>> >
>>> > Hi all,
>>> > Thanks for the engagement and support during the working group session
>>> at IETF126.
>>> > As mentioned during the meeting, we are looking for even more feedback
>>> on the Crawler Best Practices draft:
>>> https://datatracker.ietf.org/doc/draft-illyes-webbotauth-cbcp/ /
>>> https://github.com/garyillyes/cbcp Specifically, input on the scope,
>>> rate control conventions, and handling of non-malicious crawlers would be
>>> really useful to help move the draft forward. Please send any comments to
>>> the list or open issues directly on GitHub.
>>> > Thanks,
>>> > Gary
>>> > PS: if you have ideas where to send the following two drafts that CBCP
>>> depends on, that would be hugely welcome
>>> > 1. JAFAR:
>>> https://datatracker.ietf.org/doc/draft-illyes-webbotauth-jafar/
>>> > 2. REP-ext https://datatracker.ietf.org/doc/draft-illyes-repext/
>>> > _______________________________________________
>>> > Web-bot-auth mailing list -- web-bot-auth@ietf.org
>>> > To unsubscribe send an email to web-bot-auth-leave@ietf.org
>>>
>>> _______________________________________________
>>> Web-bot-auth mailing list -- web-bot-auth@ietf.org
>>> To unsubscribe send an email to web-bot-auth-leave@ietf.org
>>>
>> _______________________________________________
>> Web-bot-auth mailing list -- web-bot-auth@ietf.org
>> To unsubscribe send an email to web-bot-auth-leave@ietf.org
>>
>> _______________________________________________
> Web-bot-auth mailing list -- web-bot-auth@ietf.org
> To unsubscribe send an email to web-bot-auth-leave@ietf.org
>