[Web-bot-auth] Re: Follow-up from IETF 126: Comments on draft-illyes-webbotauth-cbcp

Shivdeep Singh <shivdeepsachdeva@gmail.com> Fri, 21 August 2026 07:17 UTC

Return-Path: <shivdeepsachdeva@gmail.com>
X-Original-To: web-bot-auth@mail2.ietf.org
Delivered-To: web-bot-auth@mail2.ietf.org
Received: from localhost (localhost [127.0.0.1]) by mail2.ietf.org (Postfix) with ESMTP id 4942412D38D78 for <web-bot-auth@mail2.ietf.org>; Fri, 21 Aug 2026 00:17:32 -0700 (PDT)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=ietf.org; s=ietf1; t=1787296652; bh=OidF0F6ojJjR3ZL8HwDboQtYNmjKdAo97fAl9YQekWs=; h=References:In-Reply-To:From:Date:Subject:To:Cc; b=wsyOEc+vY3oDLEIjVZRhC3qMtBh5NHvdKs2lXEfOe0JUxfF3wOtx9hLUehoUTYeEq qTxGXV4DQLws++csOVx/gOW3sWPiISZ93IQtaeUpretBD7yuioe2+M9hooj24Rjl77 KsS/bYWGL7ALAUk6VpPYVyP7C94c/4zJocbXulRg=
X-Virus-Scanned: amavisd-new at ietf.org
X-Spam-Flag: NO
X-Spam-Score: -2.098
X-Spam-Level:
X-Spam-Status: No, score=-2.098 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, FREEMAIL_FROM=0.001, HTML_MESSAGE=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_HELO_NONE=0.001, SPF_PASS=-0.001] autolearn=unavailable autolearn_force=no
Authentication-Results: mail2.ietf.org (amavisd-new); dkim=pass (2048-bit key) header.d=gmail.com
Received: from mail2.ietf.org ([166.84.6.31]) by localhost (mail2.ietf.org [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id gRZlSgIPVw9q for <web-bot-auth@mail2.ietf.org>; Fri, 21 Aug 2026 00:17:29 -0700 (PDT)
Received: from mail-oa1-x2d.google.com (mail-oa1-x2d.google.com [IPv6:2001:4860:4864:20::2d]) (using TLSv1.3 with cipher TLS_AES_128_GCM_SHA256 (128/128 bits) key-exchange X25519 server-signature ECDSA (P-256) server-digest SHA256) (No client certificate requested) by mail2.ietf.org (Postfix) with ESMTPS id C9ED912D38D67 for <web-bot-auth@ietf.org>; Fri, 21 Aug 2026 00:17:29 -0700 (PDT)
Received: by mail-oa1-x2d.google.com with SMTP id 586e51a60fabf-44cf70de986so444695fac.0 for <web-bot-auth@ietf.org>; Fri, 21 Aug 2026 00:17:29 -0700 (PDT)
ARC-Seal: i=1; a=rsa-sha256; t=1787296649; cv=none; d=google.com; s=arc-20260327; b=jhKWzB1Ecl4oTnIe5UkUywe5pdiz06Ib5RbxLy3iFXhF1MzbQlDPZLnDcMHMrK8B8j UmagfDmq0jTYudyzMaxI4XHxnRu3BjSwDzdUlZpY79eIvQ8E06V3EG2+mkmXg2ND7WWt 0jAJyAUqICWleqD22ZnfCAueVKfBT0BX9a4wS7LcNUAfDo1USv1q1rdyeUuE7P9n5Trl 2WMwhzek7cgWRjHDVauKGD0Gvin6gaRyrXQjv3fpGmF1Xl8mBTbdFUf9LG4TXL/vFDVd /KLVbhxxyfuzF94e5TAawaHPuHs998LNDiUMrOZbZYFvsZrbnKPEH9GGJmm6iCNr8vyj QZ6A==
ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=arc-20260327; h=cc:to:subject:message-id:date:from:in-reply-to:references :mime-version:dkim-signature; bh=0V9EULAmcd0JPqKN3AHYyfHmpOvIU/wYMHGI1LD7x7A=; fh=GqhzvlKIFn2GbUy1YH5nWbjjNR5gKtCxLUhVpwVV8Lo=; b=c/coVeB4iwxG8Y3YyfxXKKdtnoROvRMs/+fh1f20ElvT5IKv5LUBbHbzAsSRxa3Tc/ opZ3m/cJmhPK+AcTtihYyKQpLhgiKylBlwjoCXCJeJ/+/7hFCtp+NFKEHjVzU0tOjUSN IGtlE/kPab/j9CSd3z0P1r3cTrXf2BbEWayL+jTz9jy1hj8+MklpShIfTvwzl5lOmrjt aCfET5fhDvmg7XDxEZkkOjBcAp6hGAq6K+rKQ8C+DXV6qrz19ibBE59SYrYnjsHcw/ZC x2Cb1W006TjOhbzGUreZqBLZby/yc0UUibHTaKEK0Qjr/X1ZWQRY/HHkm4h3xjkG55GP ZF4g==; darn=ietf.org
ARC-Authentication-Results: i=1; mx.google.com; arc=none
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787296649; x=1787901449; darn=ietf.org; h=content-type:cc:to:subject:message-id:date:from:in-reply-to :references:mime-version:from:to:cc:subject:date:message-id:reply-to :content-type; bh=0V9EULAmcd0JPqKN3AHYyfHmpOvIU/wYMHGI1LD7x7A=; b=j2xoWNObj7deG543n9UCEnNIbO64J4m/+WHFoG5qXdiQnWUwzMvu21suuxwnV1XJq5 HN7HtKYYQFyTqc3t+TY7cnTO/23AYqOM31m+CPGEFISxIGcDF63LlbmSY9YlX47ubw0a hCamfLpH55VYpOGImcotgaFqCSBsKONEGmshsoOU1vtmz3I3ufvmzWzfVodAyhM70EE1 Bze9FPioxXVCOe1XWRCwP/E73KnYfccYsFvhnu/GaaOeFgYp6srOlY0Fha7KNlMgurOX DpX70rZdrFz4+q+W9rNLYrRHvKycJzP9xZRCLmnENCKve3A2Rxh9SxFzQDPZv35EetS9 SCIA==
X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787296649; x=1787901449; h=content-type:cc:to:subject:message-id:date:from:in-reply-to :references:mime-version:x-gm-gg:x-gm-message-state:from:to:cc :subject:date:message-id:reply-to:content-type; bh=0V9EULAmcd0JPqKN3AHYyfHmpOvIU/wYMHGI1LD7x7A=; b=Uo2a1DhIjeTcmq1HgFosdMA5dHPhHSMLK5dRglWaoIGIAysQA3Ox1qlzp0uUiNzCgV X3+bG/uQeHL5BSo0BXiZozf3iqw2+SUz1pUYp0Xke0GrfqiBljMIT28tyZrOTenZLShR luRXT5hVdKZl/IHpaTwY+B6v/xOjnbEGGmSlW9/+ONu3cGS9rsFpVRvdJ++kVfu0y8B5 hgvAqu3br3VYS0GpsmOeZmi8oSvS0kmPkKTQrDP40soxklVkfa6xg0oN+UViBijB1mWu ldRTXxnUgAmamL85Iz50afh9kfcbqIb/5Pf3JXnJAT+5zLNzCQGzm/Wpna4pdigGRiCf UTGg==
X-Forwarded-Encrypted: i=1; AHgh+RpUYt431G/Ptlp9Uuq/p2qKbzBFaseW9+DY1j2OlRGt5qbvwVSP6MBO9A+SLSW5ZJ/50p5K/PZEZGPa0dw=@ietf.org
X-Gm-Message-State: AOJu0YwLbSDvcPpEAVc7sScNrmbZyKO1O0d+a2sXg6hcdL7Uhzg4TFBl xryjA4k9AeuWmvR7wNi7b+b/BVnjp/YnDwkwqhB2CE5X2YyQ7ErbEnoq4L21jIPClsu+XvgFCZ6 A9Js0TR4sK/CZrgUJUbRCburnVEA5Cto=
X-Gm-Gg: AR+sD11+BffAlnSOmv57/v/2H7PmvUfSYUxlMTydY3z/QTmxDyNVkmak6L+LIOHWMed /s0ggNacrv5BkrRLVf+ZP04qZXQM/Po3OKghJMdEHNet74brRVgcaIxQDUUzXizAJcP7LIz1DHR XWnxawoB+SCUFhfMq1Isy3nxd2s5SYEFaX76Nu4xq3MBa39qOJY3CJ0QI8Kt7hTirV59hXuLepT j7MR0VaaEri933IqbyV8bMooHLQAuaI2+Mg6+9lvM/OoN3tgGy8dgxai4d+gLfEyPFlaVJQIWQp SlDsoBZTOCXaShNJGvLgbaUp1fkVoR+Az9XlcRHQBq6VLqZWFEtogxbXYkFVd2oukt7OMxBb4Zf JR+8=
X-Received: by 2002:a05:6820:2d05:b0:6a1:8192:4d89 with SMTP id 006d021491bc7-6b1594bd3f2mr4356360eaf.29.1787296648939; Fri, 21 Aug 2026 00:17:28 -0700 (PDT)
MIME-Version: 1.0
References: <CADTQi=dSGTGJtczLTZ1ua77C8h4pvGhrVkz3MZV7SO280MJrKQ@mail.gmail.com> <19B52E65-D016-4606-A329-70D6FF10E955@vaibhavbajpai.com> <CAKaUSfUFR82mCjy7qWzgibnuatQ7YdJ_yStur3b2zMR0KbMWkQ@mail.gmail.com> <6d65e967-b8a0-46be-9d02-00b66adb306a@Spark>
In-Reply-To: <6d65e967-b8a0-46be-9d02-00b66adb306a@Spark>
From: Shivdeep Singh <shivdeepsachdeva@gmail.com>
Date: Fri, 21 Aug 2026 12:47:15 +0530
X-Gm-Features: AcwNN1VoRDPkBpnFKb8RLc6WOiuJpIZWepsraSJYACfcxyYLIY_rKAOD47VZ0tA
Message-ID: <CAKaUSfUH=yZiHgP0=YjCnEpFKyVC3meZV9MYH+JfjYqQcF1sfA@mail.gmail.com>
To: nick@lifelightlabs.com
Content-Type: multipart/alternative; boundary="000000000000de4d8f0659896d8c"
Message-ID-Hash: EJTALP5F7CHHO6JB4E3WHN6UVZP4T34M
X-Message-ID-Hash: EJTALP5F7CHHO6JB4E3WHN6UVZP4T34M
X-MailFrom: shivdeepsachdeva@gmail.com
X-Mailman-Rule-Misses: dmarc-mitigation; no-senders; approved; emergency; loop; banned-address; member-moderation; header-match-web-bot-auth.ietf.org-0; header-match-web-bot-auth.ietf.org-1; nonmember-moderation; administrivia; implicit-dest; max-recipients; max-size; news-moderation; no-subject; digests; suspicious-header
CC: Vaibhav Bajpai <contact@vaibhavbajpai.com>, web-bot-auth@ietf.org, "Mirja Kuehlewind (IETF)" <ietf@kuehlewind.net>, Sauron <sauron=40google.com@dmarc.ietf.org>
X-Mailman-Version: 3.3.9rc6
Precedence: list
Subject: [Web-bot-auth] Re: Follow-up from IETF 126: Comments on draft-illyes-webbotauth-cbcp
List-Id: Authentication of non-human users to human-oriented Web sites <web-bot-auth.ietf.org>
Archived-At: <https://mailarchive.ietf.org/arch/msg/web-bot-auth/pXKM0dkbAWsMqcAaXObxB6WUID4>
List-Archive: <https://mailarchive.ietf.org/arch/browse/web-bot-auth>
List-Help: <mailto:web-bot-auth-request@ietf.org?subject=help>
List-Owner: <mailto:web-bot-auth-owner@ietf.org>
List-Post: <mailto:web-bot-auth@ietf.org>
List-Subscribe: <mailto:web-bot-auth-join@ietf.org>
List-Unsubscribe: <mailto:web-bot-auth-leave@ietf.org>

Nick, thank you, and the tag is the sharper version of this. I had the
scope sentence and the consumption path in 5.5.3. That a conformant proof
must carry tag=http-message-signatures-directory is what makes it concrete,
since that value has no story on a jwks_uri response.

Agreed that the ambiguity is the problem rather than either reading.

This is not hypothetical at our end. We serve a hosted directory carrying
four operators' keys under one identifier. Per-tenant authorities would
resolve that, but they require an origin per operator, so the route for
operators without one is per-tenant jwks_uri, which lands squarely in the
gap.

So if Thibault lands on directory-only, I am happy to draft the sentence
for wherever a jwks_uri or cimd operator would actually find it: that key
material published under those types cannot carry an Appendix B proof, and
that a verifier receiving it redistributed falls back to the thumbprint
identifier in 4.3, which does not survive rotation. Happy to open it as a
PR, or leave it with the editors.

Shivdeep

On Fri, Aug 21, 2026 at 8:49 AM Nick Mathews <nick@lifelightlabs.com> wrote:

> Shivdeep, Vaibhav, Gary,
>
> Speaking to the Appendix B question as the contributor of the test vectors
> that validate it (E.2.3), not for the spec text itself.
>
> I think both of your readings have support in the text as written, which
> is the problem. The framing sentence describes the concern: what a verifier
> checks when it wants the domain a key is published under, rather than the
> URL on its own. Nothing in the B.1 construction consumes a directory-format
> field; it is a per-key signature over @authority and content-digest with
> created and expires, and that applies mechanically to any HTTPS response
> that serves key material, including a hosted jwks_uri. But the normative
> hooks are directory-shaped: the required tag value is literally
> http-message-signatures-directory, and the consumption path runs through
> 5.5.3. So an implementer today cannot claim conformant proofs on a jwks_uri
> response even though the mechanism would work there.
>
> Our verifier implements B.1 on directory responses only, because that is
> what the text licenses and what the vectors cover. If the intent is
> any-type, the appendix should say so and give the tag a story for
> non-directory responses. If the intent is directory-only, then the
> consequence you describe, jwks_uri and cimd operators falling back to
> thumbprint identity with no rotation per 4.3, deserves a sentence where
> those operators will find it. Either sentence is cheap and the ambiguity is
> not. I would like to see Thibault confirm the intended scope before the BCP
> points operators at the mechanism.
>
> On the point 3 convergence, one merchant-side datapoint: distinguishable
> per-purpose identities are not just registry hygiene. Our policy layer keys
> rules on the resolved identifier, so a crawler that separates search from
> training from assistant retrieval gets separate policy rows a merchant can
> treat differently. A single generic identity collapses that to one row, and
> the measurement in the paper matches what we see from merchants: they want
> to say yes to one purpose and no to another, and the identifier is the only
> handle they have.
>
> Nick
> On Aug 20, 2026 at 6:26 AM -0400, Shivdeep Singh <
> shivdeepsachdeva@gmail.com>, wrote:
>
> Vaibhav, Gary, Nick,
>
> Point 3 and Nick's argument converge on the same mechanism, and I have a
> scope question about it.
>
> Under draft-meunier-webbotauth-httpsig-protocol-02 an identity is the
> resolved Signature-Agent URL, so distinguishable per-purpose crawlers mean
> distinct URLs. An operator without a stable origin cannot use the directory
> type, since 5.5 requires that member value to be an origin, which leaves
> jwks_uri or cimd.
>
> Appendix B opens with "It applies to the directory type in Section 5.5." I
> read that two ways. Either the possession proof in B.1 is available only to
> the directory type, in which case jwks_uri and cimd material can never
> satisfy 5.5.3 and falls back to thumbprint identity, which 4.3 says has no
> rotation. Or the scope line is about domain binding specifically, the
> concern of that appendix, and B.1's mechanism is available to any type, in
> which case saying so would help implementers.
>
> I do not think this changes either recommendation. It affects who can act
> on them, so it seemed worth asking before the BCP points operators at the
> mechanism.
>
> Shivdeep Singh
>
> On Thu, Aug 20, 2026 at 2:53 PM Vaibhav Bajpai <contact@vaibhavbajpai.com>
> wrote:
>
>> Hi everyone,
>>
>> This paper just got published, and we believe would be
>> valuable input to the BCP draft:
>>
>> >From robots.txt to ai.txt: Mapping the Evolution of Web Permissions in
>> the Age of AI
>> https://dl.acm.org/doi/pdf/10.1145/3831956.3831960
>>
>> The measurements suggest that being identifiable and respecting
>> robots.txt is
>> necessary, but no longer sufficient. The harder emerging problem is
>> ensuring
>> that authenticated crawlers interpret a fragmented set of publisher
>> signals
>> consistently, predictably, and according to the intended type of AI use.
>>
>> Our measurements show that today’s system combines heavy dependence on
>> robots.txt,
>> fragmented AI-specific policies, inconsistent naming, and only partially
>> coherent configurations.
>>
>>
>> To this end, here is some input for the BCP:
>>
>> 1.) Define how crawlers should handle conflicting signals and precedence:
>> Explicitly establish that discovery/descriptor files do not grant access
>> and
>> recommend conservative handling of contradictory permission signals.
>>
>> 2.) Separate crawler identity by purpose. Search, training, and
>> user-triggered/assistant retrieval should be distinguishable where they
>> represent materially different uses. Measurements show that that
>> publishers
>> are expressing use-specific preferences: e.g., search=yes, ai-train=no.
>>
>> 3.) Strengthen identity requirements and encourage a canonical crawler
>> registry: Documentation alone isn’t solving crawler identity, as is
>> apparent
>> from the very large number of unrecognized User-Agent strings in our
>> measurements. Plus, when an operator runs crawlers for materially
>> different
>> purposes, those crawlers should have distinguishable identities, e.g.:
>>
>>   ExampleSearchBot
>>   ExampleTrainingBot
>>   ExampleAssistantFetcher
>>
>>   rather than one generic ExampleBot.
>>
>> 4.) Explicitly prohibit identity evasion/policy shopping: A crawler
>> shouldn’t
>> spoof identities, rotate identifiers to evade rules, or choose whichever
>> applicable policy gives it the most permissive result.
>>
>> 5.) Encourage policy revalidation rather than indefinitely caching
>> permission
>> decisions: Crawlers should periodically revalidate cached crawler-policy
>> resources and must not assume that a previously observed permission
>> remains
>> valid indefinitely.
>>
>> 6.) Require transparent interpretation and validation. Crawler operators
>> should
>> document supported policy mechanisms, conflict behavior and caching, and
>> ideally provide a URL testing tool showing publishers exactly how their
>> crawler interprets the site’s policies, e.g., “This is how our crawler
>> interprets your site.”
>>
>> Kind regards, Vaibhav
>>
>> > On 27. Jul 2026, at 14:42, Sauron <sauron=40google.com@dmarc.ietf.org>
>> wrote:
>> >
>> > Hi all,
>> > Thanks for the engagement and support during the working group session
>> at IETF126.
>> > As mentioned during the meeting, we are looking for even more feedback
>> on the Crawler Best Practices draft:
>> https://datatracker.ietf.org/doc/draft-illyes-webbotauth-cbcp/ /
>> https://github.com/garyillyes/cbcp Specifically, input on the scope,
>> rate control conventions, and handling of non-malicious crawlers would be
>> really useful to help move the draft forward. Please send any comments to
>> the list or open issues directly on GitHub.
>> > Thanks,
>> > Gary
>> > PS: if you have ideas where to send the following two drafts that CBCP
>> depends on, that would be hugely welcome
>> > 1. JAFAR:
>> https://datatracker.ietf.org/doc/draft-illyes-webbotauth-jafar/
>> > 2. REP-ext https://datatracker.ietf.org/doc/draft-illyes-repext/
>> > _______________________________________________
>> > Web-bot-auth mailing list -- web-bot-auth@ietf.org
>> > To unsubscribe send an email to web-bot-auth-leave@ietf.org
>>
>> _______________________________________________
>> Web-bot-auth mailing list -- web-bot-auth@ietf.org
>> To unsubscribe send an email to web-bot-auth-leave@ietf.org
>>
> _______________________________________________
> Web-bot-auth mailing list -- web-bot-auth@ietf.org
> To unsubscribe send an email to web-bot-auth-leave@ietf.org
>
>