Return-Path: <shivdeepsachdeva@gmail.com>
X-Original-To: web-bot-auth@mail2.ietf.org
Delivered-To: web-bot-auth@mail2.ietf.org
Received: from localhost (localhost [127.0.0.1])
	by mail2.ietf.org (Postfix) with ESMTP id 4942412D38D78
	for <web-bot-auth@mail2.ietf.org>; Fri, 21 Aug 2026 00:17:32 -0700 (PDT)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=ietf.org; s=ietf1;
	t=1787296652; bh=OidF0F6ojJjR3ZL8HwDboQtYNmjKdAo97fAl9YQekWs=;
	h=References:In-Reply-To:From:Date:Subject:To:Cc;
	b=wsyOEc+vY3oDLEIjVZRhC3qMtBh5NHvdKs2lXEfOe0JUxfF3wOtx9hLUehoUTYeEq
	 qTxGXV4DQLws++csOVx/gOW3sWPiISZ93IQtaeUpretBD7yuioe2+M9hooj24Rjl77
	 KsS/bYWGL7ALAUk6VpPYVyP7C94c/4zJocbXulRg=
X-Virus-Scanned: amavisd-new at ietf.org
X-Spam-Flag: NO
X-Spam-Score: -2.098
X-Spam-Level: 
X-Spam-Status: No, score=-2.098 tagged_above=-999 required=5
	tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1,
	DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, FREEMAIL_FROM=0.001,
	HTML_MESSAGE=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_HELO_NONE=0.001,
	SPF_PASS=-0.001] autolearn=unavailable autolearn_force=no
Authentication-Results: mail2.ietf.org (amavisd-new); dkim=pass (2048-bit key)
	header.d=gmail.com
Received: from mail2.ietf.org ([166.84.6.31])
	by localhost (mail2.ietf.org [127.0.0.1]) (amavisd-new, port 10024)
	with ESMTP id gRZlSgIPVw9q for <web-bot-auth@mail2.ietf.org>;
	Fri, 21 Aug 2026 00:17:29 -0700 (PDT)
Received: from mail-oa1-x2d.google.com (mail-oa1-x2d.google.com
 [IPv6:2001:4860:4864:20::2d])
	(using TLSv1.3 with cipher TLS_AES_128_GCM_SHA256 (128/128 bits)
	 key-exchange X25519 server-signature ECDSA (P-256) server-digest SHA256)
	(No client certificate requested)
	by mail2.ietf.org (Postfix) with ESMTPS id C9ED912D38D67
	for <web-bot-auth@ietf.org>; Fri, 21 Aug 2026 00:17:29 -0700 (PDT)
Received: by mail-oa1-x2d.google.com with SMTP id
 586e51a60fabf-44cf70de986so444695fac.0
        for <web-bot-auth@ietf.org>; Fri, 21 Aug 2026 00:17:29 -0700 (PDT)
ARC-Seal: i=1; a=rsa-sha256; t=1787296649; cv=none;
        d=google.com; s=arc-20260327;
        b=jhKWzB1Ecl4oTnIe5UkUywe5pdiz06Ib5RbxLy3iFXhF1MzbQlDPZLnDcMHMrK8B8j
         UmagfDmq0jTYudyzMaxI4XHxnRu3BjSwDzdUlZpY79eIvQ8E06V3EG2+mkmXg2ND7WWt
         0jAJyAUqICWleqD22ZnfCAueVKfBT0BX9a4wS7LcNUAfDo1USv1q1rdyeUuE7P9n5Trl
         2WMwhzek7cgWRjHDVauKGD0Gvin6gaRyrXQjv3fpGmF1Xl8mBTbdFUf9LG4TXL/vFDVd
         /KLVbhxxyfuzF94e5TAawaHPuHs998LNDiUMrOZbZYFvsZrbnKPEH9GGJmm6iCNr8vyj
         QZ6A==
ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com;
 s=arc-20260327;
        h=cc:to:subject:message-id:date:from:in-reply-to:references
         :mime-version:dkim-signature;
        bh=0V9EULAmcd0JPqKN3AHYyfHmpOvIU/wYMHGI1LD7x7A=;
        fh=GqhzvlKIFn2GbUy1YH5nWbjjNR5gKtCxLUhVpwVV8Lo=;
        b=c/coVeB4iwxG8Y3YyfxXKKdtnoROvRMs/+fh1f20ElvT5IKv5LUBbHbzAsSRxa3Tc/
         opZ3m/cJmhPK+AcTtihYyKQpLhgiKylBlwjoCXCJeJ/+/7hFCtp+NFKEHjVzU0tOjUSN
         IGtlE/kPab/j9CSd3z0P1r3cTrXf2BbEWayL+jTz9jy1hj8+MklpShIfTvwzl5lOmrjt
         aCfET5fhDvmg7XDxEZkkOjBcAp6hGAq6K+rKQ8C+DXV6qrz19ibBE59SYrYnjsHcw/ZC
         x2Cb1W006TjOhbzGUreZqBLZby/yc0UUibHTaKEK0Qjr/X1ZWQRY/HHkm4h3xjkG55GP
         ZF4g==;
        darn=ietf.org
ARC-Authentication-Results: i=1; mx.google.com; arc=none
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed;
        d=gmail.com; s=20251104; t=1787296649; x=1787901449; darn=ietf.org;
        h=content-type:cc:to:subject:message-id:date:from:in-reply-to
         :references:mime-version:from:to:cc:subject:date:message-id:reply-to
         :content-type;
        bh=0V9EULAmcd0JPqKN3AHYyfHmpOvIU/wYMHGI1LD7x7A=;
        b=j2xoWNObj7deG543n9UCEnNIbO64J4m/+WHFoG5qXdiQnWUwzMvu21suuxwnV1XJq5
         HN7HtKYYQFyTqc3t+TY7cnTO/23AYqOM31m+CPGEFISxIGcDF63LlbmSY9YlX47ubw0a
         hCamfLpH55VYpOGImcotgaFqCSBsKONEGmshsoOU1vtmz3I3ufvmzWzfVodAyhM70EE1
         Bze9FPioxXVCOe1XWRCwP/E73KnYfccYsFvhnu/GaaOeFgYp6srOlY0Fha7KNlMgurOX
         DpX70rZdrFz4+q+W9rNLYrRHvKycJzP9xZRCLmnENCKve3A2Rxh9SxFzQDPZv35EetS9
         SCIA==
X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed;
        d=1e100.net; s=20251104; t=1787296649; x=1787901449;
        h=content-type:cc:to:subject:message-id:date:from:in-reply-to
         :references:mime-version:x-gm-gg:x-gm-message-state:from:to:cc
         :subject:date:message-id:reply-to:content-type;
        bh=0V9EULAmcd0JPqKN3AHYyfHmpOvIU/wYMHGI1LD7x7A=;
        b=Uo2a1DhIjeTcmq1HgFosdMA5dHPhHSMLK5dRglWaoIGIAysQA3Ox1qlzp0uUiNzCgV
         X3+bG/uQeHL5BSo0BXiZozf3iqw2+SUz1pUYp0Xke0GrfqiBljMIT28tyZrOTenZLShR
         luRXT5hVdKZl/IHpaTwY+B6v/xOjnbEGGmSlW9/+ONu3cGS9rsFpVRvdJ++kVfu0y8B5
         hgvAqu3br3VYS0GpsmOeZmi8oSvS0kmPkKTQrDP40soxklVkfa6xg0oN+UViBijB1mWu
         ldRTXxnUgAmamL85Iz50afh9kfcbqIb/5Pf3JXnJAT+5zLNzCQGzm/Wpna4pdigGRiCf
         UTGg==
X-Forwarded-Encrypted: i=1;
 AHgh+RpUYt431G/Ptlp9Uuq/p2qKbzBFaseW9+DY1j2OlRGt5qbvwVSP6MBO9A+SLSW5ZJ/50p5K/PZEZGPa0dw=@ietf.org
X-Gm-Message-State: AOJu0YwLbSDvcPpEAVc7sScNrmbZyKO1O0d+a2sXg6hcdL7Uhzg4TFBl
	xryjA4k9AeuWmvR7wNi7b+b/BVnjp/YnDwkwqhB2CE5X2YyQ7ErbEnoq4L21jIPClsu+XvgFCZ6
	A9Js0TR4sK/CZrgUJUbRCburnVEA5Cto=
X-Gm-Gg: AR+sD11+BffAlnSOmv57/v/2H7PmvUfSYUxlMTydY3z/QTmxDyNVkmak6L+LIOHWMed
	/s0ggNacrv5BkrRLVf+ZP04qZXQM/Po3OKghJMdEHNet74brRVgcaIxQDUUzXizAJcP7LIz1DHR
	XWnxawoB+SCUFhfMq1Isy3nxd2s5SYEFaX76Nu4xq3MBa39qOJY3CJ0QI8Kt7hTirV59hXuLepT
	j7MR0VaaEri933IqbyV8bMooHLQAuaI2+Mg6+9lvM/OoN3tgGy8dgxai4d+gLfEyPFlaVJQIWQp
	SlDsoBZTOCXaShNJGvLgbaUp1fkVoR+Az9XlcRHQBq6VLqZWFEtogxbXYkFVd2oukt7OMxBb4Zf
	JR+8=
X-Received: by 2002:a05:6820:2d05:b0:6a1:8192:4d89 with SMTP id
 006d021491bc7-6b1594bd3f2mr4356360eaf.29.1787296648939; Fri, 21 Aug 2026
 00:17:28 -0700 (PDT)
MIME-Version: 1.0
References: 
 <CADTQi=dSGTGJtczLTZ1ua77C8h4pvGhrVkz3MZV7SO280MJrKQ@mail.gmail.com>
 <19B52E65-D016-4606-A329-70D6FF10E955@vaibhavbajpai.com>
 <CAKaUSfUFR82mCjy7qWzgibnuatQ7YdJ_yStur3b2zMR0KbMWkQ@mail.gmail.com>
 <6d65e967-b8a0-46be-9d02-00b66adb306a@Spark>
In-Reply-To: <6d65e967-b8a0-46be-9d02-00b66adb306a@Spark>
From: Shivdeep Singh <shivdeepsachdeva@gmail.com>
Date: Fri, 21 Aug 2026 12:47:15 +0530
X-Gm-Features: AcwNN1VoRDPkBpnFKb8RLc6WOiuJpIZWepsraSJYACfcxyYLIY_rKAOD47VZ0tA
Message-ID: 
 <CAKaUSfUH=yZiHgP0=YjCnEpFKyVC3meZV9MYH+JfjYqQcF1sfA@mail.gmail.com>
To: nick@lifelightlabs.com
Content-Type: multipart/alternative; boundary="000000000000de4d8f0659896d8c"
Message-ID-Hash: EJTALP5F7CHHO6JB4E3WHN6UVZP4T34M
X-Message-ID-Hash: EJTALP5F7CHHO6JB4E3WHN6UVZP4T34M
X-MailFrom: shivdeepsachdeva@gmail.com
X-Mailman-Rule-Misses: dmarc-mitigation; no-senders; approved; emergency;
 loop; banned-address; member-moderation;
 header-match-web-bot-auth.ietf.org-0; header-match-web-bot-auth.ietf.org-1;
 nonmember-moderation; administrivia; implicit-dest; max-recipients; max-size;
 news-moderation; no-subject; digests; suspicious-header
CC: Vaibhav Bajpai <contact@vaibhavbajpai.com>, web-bot-auth@ietf.org,
 "Mirja Kuehlewind (IETF)" <ietf@kuehlewind.net>,
 Sauron <sauron=40google.com@dmarc.ietf.org>
X-Mailman-Version: 3.3.9rc6
Precedence: list
Subject: =?utf-8?q?=5BWeb-bot-auth=5D_Re=3A_Follow-up_from_IETF_126=3A_Comments_on_dr?=
	=?utf-8?q?aft-illyes-webbotauth-cbcp?=
List-Id: Authentication of non-human users to human-oriented Web sites
 <web-bot-auth.ietf.org>
Archived-At: 
 <https://mailarchive.ietf.org/arch/msg/web-bot-auth/pXKM0dkbAWsMqcAaXObxB6WUID4>
List-Archive: <https://mailarchive.ietf.org/arch/browse/web-bot-auth>
List-Help: <mailto:web-bot-auth-request@ietf.org?subject=help>
List-Owner: <mailto:web-bot-auth-owner@ietf.org>
List-Post: <mailto:web-bot-auth@ietf.org>
List-Subscribe: <mailto:web-bot-auth-join@ietf.org>
List-Unsubscribe: <mailto:web-bot-auth-leave@ietf.org>

--000000000000de4d8f0659896d8c
Content-Type: text/plain; charset="UTF-8"
Content-Transfer-Encoding: quoted-printable

Nick, thank you, and the tag is the sharper version of this. I had the
scope sentence and the consumption path in 5.5.3. That a conformant proof
must carry tag=3Dhttp-message-signatures-directory is what makes it concret=
e,
since that value has no story on a jwks_uri response.

Agreed that the ambiguity is the problem rather than either reading.

This is not hypothetical at our end. We serve a hosted directory carrying
four operators' keys under one identifier. Per-tenant authorities would
resolve that, but they require an origin per operator, so the route for
operators without one is per-tenant jwks_uri, which lands squarely in the
gap.

So if Thibault lands on directory-only, I am happy to draft the sentence
for wherever a jwks_uri or cimd operator would actually find it: that key
material published under those types cannot carry an Appendix B proof, and
that a verifier receiving it redistributed falls back to the thumbprint
identifier in 4.3, which does not survive rotation. Happy to open it as a
PR, or leave it with the editors.

Shivdeep

On Fri, Aug 21, 2026 at 8:49=E2=80=AFAM Nick Mathews <nick@lifelightlabs.co=
m> wrote:

> Shivdeep, Vaibhav, Gary,
>
> Speaking to the Appendix B question as the contributor of the test vector=
s
> that validate it (E.2.3), not for the spec text itself.
>
> I think both of your readings have support in the text as written, which
> is the problem. The framing sentence describes the concern: what a verifi=
er
> checks when it wants the domain a key is published under, rather than the
> URL on its own. Nothing in the B.1 construction consumes a directory-form=
at
> field; it is a per-key signature over @authority and content-digest with
> created and expires, and that applies mechanically to any HTTPS response
> that serves key material, including a hosted jwks_uri. But the normative
> hooks are directory-shaped: the required tag value is literally
> http-message-signatures-directory, and the consumption path runs through
> 5.5.3. So an implementer today cannot claim conformant proofs on a jwks_u=
ri
> response even though the mechanism would work there.
>
> Our verifier implements B.1 on directory responses only, because that is
> what the text licenses and what the vectors cover. If the intent is
> any-type, the appendix should say so and give the tag a story for
> non-directory responses. If the intent is directory-only, then the
> consequence you describe, jwks_uri and cimd operators falling back to
> thumbprint identity with no rotation per 4.3, deserves a sentence where
> those operators will find it. Either sentence is cheap and the ambiguity =
is
> not. I would like to see Thibault confirm the intended scope before the B=
CP
> points operators at the mechanism.
>
> On the point 3 convergence, one merchant-side datapoint: distinguishable
> per-purpose identities are not just registry hygiene. Our policy layer ke=
ys
> rules on the resolved identifier, so a crawler that separates search from
> training from assistant retrieval gets separate policy rows a merchant ca=
n
> treat differently. A single generic identity collapses that to one row, a=
nd
> the measurement in the paper matches what we see from merchants: they wan=
t
> to say yes to one purpose and no to another, and the identifier is the on=
ly
> handle they have.
>
> Nick
> On Aug 20, 2026 at 6:26=E2=80=AFAM -0400, Shivdeep Singh <
> shivdeepsachdeva@gmail.com>, wrote:
>
> Vaibhav, Gary, Nick,
>
> Point 3 and Nick's argument converge on the same mechanism, and I have a
> scope question about it.
>
> Under draft-meunier-webbotauth-httpsig-protocol-02 an identity is the
> resolved Signature-Agent URL, so distinguishable per-purpose crawlers mea=
n
> distinct URLs. An operator without a stable origin cannot use the directo=
ry
> type, since 5.5 requires that member value to be an origin, which leaves
> jwks_uri or cimd.
>
> Appendix B opens with "It applies to the directory type in Section 5.5." =
I
> read that two ways. Either the possession proof in B.1 is available only =
to
> the directory type, in which case jwks_uri and cimd material can never
> satisfy 5.5.3 and falls back to thumbprint identity, which 4.3 says has n=
o
> rotation. Or the scope line is about domain binding specifically, the
> concern of that appendix, and B.1's mechanism is available to any type, i=
n
> which case saying so would help implementers.
>
> I do not think this changes either recommendation. It affects who can act
> on them, so it seemed worth asking before the BCP points operators at the
> mechanism.
>
> Shivdeep Singh
>
> On Thu, Aug 20, 2026 at 2:53=E2=80=AFPM Vaibhav Bajpai <contact@vaibhavba=
jpai.com>
> wrote:
>
>> Hi everyone,
>>
>> This paper just got published, and we believe would be
>> valuable input to the BCP draft:
>>
>> >From robots.txt to ai.txt: Mapping the Evolution of Web Permissions in
>> the Age of AI
>> https://dl.acm.org/doi/pdf/10.1145/3831956.3831960
>>
>> The measurements suggest that being identifiable and respecting
>> robots.txt is
>> necessary, but no longer sufficient. The harder emerging problem is
>> ensuring
>> that authenticated crawlers interpret a fragmented set of publisher
>> signals
>> consistently, predictably, and according to the intended type of AI use.
>>
>> Our measurements show that today=E2=80=99s system combines heavy depende=
nce on
>> robots.txt,
>> fragmented AI-specific policies, inconsistent naming, and only partially
>> coherent configurations.
>>
>>
>> To this end, here is some input for the BCP:
>>
>> 1.) Define how crawlers should handle conflicting signals and precedence=
:
>> Explicitly establish that discovery/descriptor files do not grant access
>> and
>> recommend conservative handling of contradictory permission signals.
>>
>> 2.) Separate crawler identity by purpose. Search, training, and
>> user-triggered/assistant retrieval should be distinguishable where they
>> represent materially different uses. Measurements show that that
>> publishers
>> are expressing use-specific preferences: e.g., search=3Dyes, ai-train=3D=
no.
>>
>> 3.) Strengthen identity requirements and encourage a canonical crawler
>> registry: Documentation alone isn=E2=80=99t solving crawler identity, as=
 is
>> apparent
>> from the very large number of unrecognized User-Agent strings in our
>> measurements. Plus, when an operator runs crawlers for materially
>> different
>> purposes, those crawlers should have distinguishable identities, e.g.:
>>
>>   ExampleSearchBot
>>   ExampleTrainingBot
>>   ExampleAssistantFetcher
>>
>>   rather than one generic ExampleBot.
>>
>> 4.) Explicitly prohibit identity evasion/policy shopping: A crawler
>> shouldn=E2=80=99t
>> spoof identities, rotate identifiers to evade rules, or choose whichever
>> applicable policy gives it the most permissive result.
>>
>> 5.) Encourage policy revalidation rather than indefinitely caching
>> permission
>> decisions: Crawlers should periodically revalidate cached crawler-policy
>> resources and must not assume that a previously observed permission
>> remains
>> valid indefinitely.
>>
>> 6.) Require transparent interpretation and validation. Crawler operators
>> should
>> document supported policy mechanisms, conflict behavior and caching, and
>> ideally provide a URL testing tool showing publishers exactly how their
>> crawler interprets the site=E2=80=99s policies, e.g., =E2=80=9CThis is h=
ow our crawler
>> interprets your site.=E2=80=9D
>>
>> Kind regards, Vaibhav
>>
>> > On 27. Jul 2026, at 14:42, Sauron <sauron=3D40google.com@dmarc.ietf.or=
g>
>> wrote:
>> >
>> > Hi all,
>> > Thanks for the engagement and support during the working group session
>> at IETF126.
>> > As mentioned during the meeting, we are looking for even more feedback
>> on the Crawler Best Practices draft:
>> https://datatracker.ietf.org/doc/draft-illyes-webbotauth-cbcp/ /
>> https://github.com/garyillyes/cbcp Specifically, input on the scope,
>> rate control conventions, and handling of non-malicious crawlers would b=
e
>> really useful to help move the draft forward. Please send any comments t=
o
>> the list or open issues directly on GitHub.
>> > Thanks,
>> > Gary
>> > PS: if you have ideas where to send the following two drafts that CBCP
>> depends on, that would be hugely welcome
>> > 1. JAFAR:
>> https://datatracker.ietf.org/doc/draft-illyes-webbotauth-jafar/
>> > 2. REP-ext https://datatracker.ietf.org/doc/draft-illyes-repext/
>> > _______________________________________________
>> > Web-bot-auth mailing list -- web-bot-auth@ietf.org
>> > To unsubscribe send an email to web-bot-auth-leave@ietf.org
>>
>> _______________________________________________
>> Web-bot-auth mailing list -- web-bot-auth@ietf.org
>> To unsubscribe send an email to web-bot-auth-leave@ietf.org
>>
> _______________________________________________
> Web-bot-auth mailing list -- web-bot-auth@ietf.org
> To unsubscribe send an email to web-bot-auth-leave@ietf.org
>
>

--000000000000de4d8f0659896d8c
Content-Type: text/html; charset="UTF-8"
Content-Transfer-Encoding: quoted-printable

<div dir=3D"ltr"><p dir=3D"ltr">Nick, thank you, and the tag is the sharper=
 version of this. I had the scope sentence and the consumption path in 5.5.=
3. That a conformant proof must carry tag=3Dhttp-message-signatures-directo=
ry is what makes it concrete, since that value has no story on a jwks_uri r=
esponse.</p><p dir=3D"ltr">Agreed that the ambiguity is the problem rather =
than either reading.</p><p dir=3D"ltr">This is not hypothetical at our end.=
 We serve a hosted directory carrying four operators&#39; keys under one id=
entifier. Per-tenant authorities would resolve that, but they require an or=
igin per operator, so the route for operators without one is per-tenant jwk=
s_uri, which lands squarely in the gap.</p><p dir=3D"ltr">So if Thibault la=
nds on directory-only, I am happy to draft the sentence for wherever a jwks=
_uri or cimd operator would actually find it: that key material published u=
nder those types cannot carry an Appendix B proof, and that a verifier rece=
iving it redistributed falls back to the thumbprint identifier in 4.3, whic=
h does not survive rotation. Happy to open it as a PR, or leave it with the=
 editors.</p><p dir=3D"ltr">Shivdeep</p></div><br><div class=3D"gmail_quote=
 gmail_quote_container"><div dir=3D"ltr" class=3D"gmail_attr">On Fri, Aug 2=
1, 2026 at 8:49=E2=80=AFAM Nick Mathews &lt;<a href=3D"mailto:nick@lifeligh=
tlabs.com">nick@lifelightlabs.com</a>&gt; wrote:<br></div><blockquote class=
=3D"gmail_quote" style=3D"margin:0px 0px 0px 0.8ex;border-left:1px solid rg=
b(204,204,204);padding-left:1ex">



<div>
<div name=3D"messageBodySection">
<div>Shivdeep, Vaibhav, Gary,</div>
<div>=C2=A0</div>
<div>Speaking to the Appendix B question as the contributor of the test vec=
tors that validate it (E.2.3), not for the spec text itself.</div>
<div>=C2=A0</div>
<div>I think both of your readings have support in the text as written, whi=
ch is the problem. The framing sentence describes the concern: what a verif=
ier checks when it wants the domain a key is published under, rather than t=
he URL on its own. Nothing in the B.1 construction consumes a directory-for=
mat field; it is a per-key signature over @authority and content-digest wit=
h created and expires, and that applies mechanically to any HTTPS response =
that serves key material, including a hosted jwks_uri. But the normative ho=
oks are directory-shaped: the required tag value is literally http-message-=
signatures-directory, and the consumption path runs through 5.5.3. So an im=
plementer today cannot claim conformant proofs on a jwks_uri response even =
though the mechanism would work there.</div>
<div>=C2=A0</div>
<div>Our verifier implements B.1 on directory responses only, because that =
is what the text licenses and what the vectors cover. If the intent is any-=
type, the appendix should say so and give the tag a story for non-directory=
 responses. If the intent is directory-only, then the consequence you descr=
ibe, jwks_uri and cimd operators falling back to thumbprint identity with n=
o rotation per 4.3, deserves a sentence where those operators will find it.=
 Either sentence is cheap and the ambiguity is not. I would like to see Thi=
bault confirm the intended scope before the BCP points operators at the mec=
hanism.</div>
<div>=C2=A0</div>
<div>On the point 3 convergence, one merchant-side datapoint: distinguishab=
le per-purpose identities are not just registry hygiene. Our policy layer k=
eys rules on the resolved identifier, so a crawler that separates search fr=
om training from assistant retrieval gets separate policy rows a merchant c=
an treat differently. A single generic identity collapses that to one row, =
and the measurement in the paper matches what we see from merchants: they w=
ant to say yes to one purpose and no to another, and the identifier is the =
only handle they have.</div>
<div>=C2=A0</div>
<div>Nick</div>
</div>
<div name=3D"messageReplySection">On Aug 20, 2026 at 6:26=E2=80=AFAM -0400,=
 Shivdeep Singh &lt;<a href=3D"mailto:shivdeepsachdeva@gmail.com" target=3D=
"_blank">shivdeepsachdeva@gmail.com</a>&gt;, wrote:<br>
<blockquote type=3D"cite">
<div dir=3D"ltr">Vaibhav, Gary, Nick,<br>
<br>
Point 3 and Nick&#39;s argument converge on the same mechanism, and I have =
a scope question about it.<br>
<br>
Under draft-meunier-webbotauth-httpsig-protocol-02 an identity is the resol=
ved Signature-Agent URL, so distinguishable per-purpose crawlers mean disti=
nct URLs. An operator without a stable origin cannot use the directory type=
, since 5.5 requires that member value to be an origin, which leaves jwks_u=
ri or cimd.<br>
<br>
Appendix B opens with &quot;It applies to the directory type in Section 5.5=
.&quot; I read that two ways. Either the possession proof in B.1 is availab=
le only to the directory type, in which case jwks_uri and cimd material can=
 never satisfy 5.5.3 and falls back to thumbprint identity, which 4.3 says =
has no rotation. Or the scope line is about domain binding specifically, th=
e concern of that appendix, and B.1&#39;s mechanism is available to any typ=
e, in which case saying so would help implementers.<br>
<br>
I do not think this changes either recommendation. It affects who can act o=
n them, so it seemed worth asking before the BCP points operators at the me=
chanism.<br>
<br>
Shivdeep Singh</div>
<br>
<div class=3D"gmail_quote">
<div dir=3D"ltr" class=3D"gmail_attr">On Thu, Aug 20, 2026 at 2:53=E2=80=AF=
PM Vaibhav Bajpai &lt;<a href=3D"mailto:contact@vaibhavbajpai.com" target=
=3D"_blank">contact@vaibhavbajpai.com</a>&gt; wrote:<br></div>
<blockquote class=3D"gmail_quote" style=3D"margin:0px 0px 0px 0.8ex;border-=
left:1px solid rgb(204,204,204);padding-left:1ex">Hi everyone,<br>
<br>
This paper just got published, and we believe would be<br>
valuable input to the BCP draft:<br>
<br>
&gt;From robots.txt to ai.txt: Mapping the Evolution of Web Permissions in =
the Age of AI<br>
<a href=3D"https://dl.acm.org/doi/pdf/10.1145/3831956.3831960" rel=3D"noref=
errer" target=3D"_blank">https://dl.acm.org/doi/pdf/10.1145/3831956.3831960=
</a><br>
<br>
The measurements suggest that being identifiable and respecting robots.txt =
is<br>
necessary, but no longer sufficient. The harder emerging problem is ensurin=
g<br>
that authenticated crawlers interpret a fragmented set of publisher signals=
<br>
consistently, predictably, and according to the intended type of AI use.<br=
>
<br>
Our measurements show that today=E2=80=99s system combines heavy dependence=
 on robots.txt,<br>
fragmented AI-specific policies, inconsistent naming, and only partially<br=
>
coherent configurations.<br>
<br>
<br>
To this end, here is some input for the BCP:<br>
<br>
1.) Define how crawlers should handle conflicting signals and precedence:<b=
r>
Explicitly establish that discovery/descriptor files do not grant access an=
d<br>
recommend conservative handling of contradictory permission signals.<br>
<br>
2.) Separate crawler identity by purpose. Search, training, and<br>
user-triggered/assistant retrieval should be distinguishable where they<br>
represent materially different uses. Measurements show that that publishers=
<br>
are expressing use-specific preferences: e.g., search=3Dyes, ai-train=3Dno.=
<br>
<br>
3.) Strengthen identity requirements and encourage a canonical crawler<br>
registry: Documentation alone isn=E2=80=99t solving crawler identity, as is=
 apparent<br>
from the very large number of unrecognized User-Agent strings in our<br>
measurements. Plus, when an operator runs crawlers for materially different=
<br>
purposes, those crawlers should have distinguishable identities, e.g.:<br>
<br>
=C2=A0 ExampleSearchBot<br>
=C2=A0 ExampleTrainingBot<br>
=C2=A0 ExampleAssistantFetcher<br>
<br>
=C2=A0 rather than one generic ExampleBot.<br>
<br>
4.) Explicitly prohibit identity evasion/policy shopping: A crawler shouldn=
=E2=80=99t<br>
spoof identities, rotate identifiers to evade rules, or choose whichever<br=
>
applicable policy gives it the most permissive result.<br>
<br>
5.) Encourage policy revalidation rather than indefinitely caching permissi=
on<br>
decisions: Crawlers should periodically revalidate cached crawler-policy<br=
>
resources and must not assume that a previously observed permission remains=
<br>
valid indefinitely.<br>
<br>
6.) Require transparent interpretation and validation. Crawler operators sh=
ould<br>
document supported policy mechanisms, conflict behavior and caching, and<br=
>
ideally provide a URL testing tool showing publishers exactly how their<br>
crawler interprets the site=E2=80=99s policies, e.g., =E2=80=9CThis is how =
our crawler<br>
interprets your site.=E2=80=9D<br>
<br>
Kind regards, Vaibhav<br>
<br>
&gt; On 27. Jul 2026, at 14:42, Sauron &lt;sauron=3D<a href=3D"mailto:40goo=
gle.com@dmarc.ietf.org" target=3D"_blank">40google.com@dmarc.ietf.org</a>&g=
t; wrote:<br>
&gt;<br>
&gt; Hi all,<br>
&gt; Thanks for the engagement and support during the working group session=
 at IETF126.<br>
&gt; As mentioned during the meeting, we are looking for even more feedback=
 on the Crawler Best Practices draft: <a href=3D"https://datatracker.ietf.o=
rg/doc/draft-illyes-webbotauth-cbcp/" rel=3D"noreferrer" target=3D"_blank">=
https://datatracker.ietf.org/doc/draft-illyes-webbotauth-cbcp/</a> / <a hre=
f=3D"https://github.com/garyillyes/cbcp" rel=3D"noreferrer" target=3D"_blan=
k">https://github.com/garyillyes/cbcp</a> Specifically, input on the scope,=
 rate control conventions, and handling of non-malicious crawlers would be =
really useful to help move the draft forward. Please send any comments to t=
he list or open issues directly on GitHub.<br>
&gt; Thanks,<br>
&gt; Gary<br>
&gt; PS: if you have ideas where to send the following two drafts that CBCP=
 depends on, that would be hugely welcome<br>
&gt; 1. JAFAR: <a href=3D"https://datatracker.ietf.org/doc/draft-illyes-web=
botauth-jafar/" rel=3D"noreferrer" target=3D"_blank">https://datatracker.ie=
tf.org/doc/draft-illyes-webbotauth-jafar/</a><br>
&gt; 2. REP-ext <a href=3D"https://datatracker.ietf.org/doc/draft-illyes-re=
pext/" rel=3D"noreferrer" target=3D"_blank">https://datatracker.ietf.org/do=
c/draft-illyes-repext/</a><br>
&gt; _______________________________________________<br>
&gt; Web-bot-auth mailing list -- <a href=3D"mailto:web-bot-auth@ietf.org" =
target=3D"_blank">web-bot-auth@ietf.org</a><br>
&gt; To unsubscribe send an email to <a href=3D"mailto:web-bot-auth-leave@i=
etf.org" target=3D"_blank">web-bot-auth-leave@ietf.org</a><br>
<br>
_______________________________________________<br>
Web-bot-auth mailing list -- <a href=3D"mailto:web-bot-auth@ietf.org" targe=
t=3D"_blank">web-bot-auth@ietf.org</a><br>
To unsubscribe send an email to <a href=3D"mailto:web-bot-auth-leave@ietf.o=
rg" target=3D"_blank">web-bot-auth-leave@ietf.org</a><br></blockquote>
</div>
_______________________________________________<br>
Web-bot-auth mailing list -- <a href=3D"mailto:web-bot-auth@ietf.org" targe=
t=3D"_blank">web-bot-auth@ietf.org</a><br>
To unsubscribe send an email to <a href=3D"mailto:web-bot-auth-leave@ietf.o=
rg" target=3D"_blank">web-bot-auth-leave@ietf.org</a><br></blockquote>
</div>
</div>

</blockquote></div>

--000000000000de4d8f0659896d8c--

