[Web-bot-auth] Re: Follow-up from IETF 126: Comments on draft-illyes-webbotauth-cbcp
Nick Mathews <nick@lifelightlabs.com> Thu, 06 August 2026 15:29 UTC
Return-Path: <nick@lifelightlabs.com>
X-Original-To: web-bot-auth@mail2.ietf.org
Delivered-To: web-bot-auth@mail2.ietf.org
Received: from localhost (localhost [127.0.0.1]) by mail2.ietf.org (Postfix) with ESMTP id 573E8124CE5B7 for <web-bot-auth@mail2.ietf.org>; Thu, 6 Aug 2026 08:29:41 -0700 (PDT)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=ietf.org; s=ietf1; t=1786030181; bh=PWdpA2wuxiDJZCdgTkKz99VNs/v0+3uYlmROje3J6xo=; h=Date:From:To:Cc:In-Reply-To:References:Subject; b=PC2yu7ROhH+1NcjvIcrPkSRCX7av+6qJmb1VV0n6VTdlMcCS2Sj9M0n7x7aFxtKQR eIQ6jyRgT4JRqEqb7kFWKo/XzWm91xW0QPdIKvOxUtLem7qcRoaHqyxtuULPFhMovs Hrxs4wwE1cWfoYDxFJHMWrtdqYP55XN4OPMZEYbI=
X-Virus-Scanned: amavisd-new at ietf.org
X-Spam-Flag: NO
X-Spam-Score: -1.689
X-Spam-Level:
X-Spam-Status: No, score=-1.689 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_INVALID=0.1, DKIM_SIGNED=0.1, HTML_MESSAGE=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_HELO_NONE=0.001, SPF_PASS=-0.001, T_KAM_HTML_FONT_INVALID=0.01] autolearn=no autolearn_force=no
Authentication-Results: mail2.ietf.org (amavisd-new); dkim=neutral reason="invalid (public key: not available)" header.d=lifelightlabs.com
Received: from mail2.ietf.org ([166.84.6.31]) by localhost (mail2.ietf.org [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id djUHQknAkRDr for <web-bot-auth@mail2.ietf.org>; Thu, 6 Aug 2026 08:29:37 -0700 (PDT)
Received: from mail-yx1-xb134.google.com (mail-yx1-xb134.google.com [IPv6:2607:f8b0:4864:20::b134]) (using TLSv1.3 with cipher TLS_AES_128_GCM_SHA256 (128/128 bits) key-exchange X25519 server-signature ECDSA (P-256) server-digest SHA256) (No client certificate requested) by mail2.ietf.org (Postfix) with ESMTPS id A5B3B124CE5AD for <web-bot-auth@ietf.org>; Thu, 6 Aug 2026 08:29:37 -0700 (PDT)
Received: by mail-yx1-xb134.google.com with SMTP id 956f58d0204a3-667971437d6so3153496d50.2 for <web-bot-auth@ietf.org>; Thu, 06 Aug 2026 08:29:37 -0700 (PDT)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=lifelightlabs.com; s=google; t=1786030171; x=1786634971; darn=ietf.org; h=content-type:mime-version:subject:references:in-reply-to:message-id :cc:to:from:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=PWdpA2wuxiDJZCdgTkKz99VNs/v0+3uYlmROje3J6xo=; b=LsgP0TdkZlsK2a8O2GX0+2XrE6eXMdjTINHSLuCmLatBt6kFR+qGpUfDapxRUxrtmq f9JfHYBkTZ3Dre++KvS3f5sqzDHNI5ebvaBknSdTqS2rCPuq5sTYxgjUOpaJXEnL6D/4 JLJLhENbW9Ik/o5lt7dAfpgdB7iB7xdD7pH2E=
X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786030171; x=1786634971; h=content-type:mime-version:subject:references:in-reply-to:message-id :cc:to:from:date:x-gm-gg:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=PWdpA2wuxiDJZCdgTkKz99VNs/v0+3uYlmROje3J6xo=; b=tStuIwNmUKytY/71ufV7vYSaL5lIsRJz5pzA7SRp+SA4p95PwS2qc0MaGLcYPIzMrg hOKYUY13chHLxIsAxJ6S0c4/MEwIOWQrHAqHyUWttxt07DW6FzxAmPIqkfmuoE3VsmUv dYYCG5nsleR4SuTMvnF4Y+hrfnd38yoB2etbX4V3Sjhz5DHpcd8bEYqCXSUESOUgZHzq wlXx5u8zASpBk/eyejMrC9gPKw3gXOFnYKFrshotQqzrxpSIAv6nR6xiG9csZ7WnDykY a5E+5GgX+MVyAZEGllJ35yABreXUMdpjI18P0GpInRL1LSseGVQ5ZsF9ibzdIyhQXiIx x3Cg==
X-Forwarded-Encrypted: i=1; AHgh+RoA9sGo5wqn0l/Hbp3lSXehtzvFrnO8EaNt/3nbv3nVpTmlq5te/smyERWMwkRWaeWSqYBtszJBMmSRrQU=@ietf.org
X-Gm-Message-State: AOJu0Yz8xHyiO33nNySNJ860noe+WL6+PTZlC+2kWIe7JwGKLSPbWMh3 EFf/a4YmbcbhSMpzgcIgtPB0nx/PzUaF68TqV2FohybTPxaLQFQ5SAcMBSZMMP75sOc=
X-Gm-Gg: AR+sD11sovHHsH1V/zZYpsWXqG2tZfevwj05pXF68/1Y0c1I9NXDU4tO8t8gcX8OKgU rwtMBE0iedXDwGpGNv7sjfnb9BYQwjMD9bt9IaZhPSATrnXvVJtgVn7ivuM9mQw6W7VborqESz2 j4Ig5HsPklUYDsEffwWqpldNydLruSe7qO9DVQMhyfsftco1rnUcZ6za0pZV4Q9vfGDpAWiPZkc d755HtzI7T7rmDQAXBagMuAG4ql1QzZessMYKYRiEnNbujKRnxt05LFATbH/ckyVQ38fDxKwEKs XwxDt+mvPzwMK3DIb07bw5bM/6rFHZeuD6LlCONrtoN6lwWkdSG7LNIA8qT/KiBywvsBYn+jyeG lz8wGW9BdfX/T/nB+fNlhw3ODOinVr9wgsvb1iy5BfuVmOm6IvBRAaNiJatj/nv5/9nxXIptq+3 jPk2QnxgBFUahymFwV00VR/igB1aEkx47Co73WTclDJlN2kkGPfQXZFiAUybiBDdTkSMtK+5BMQ 0e3SUMoPvjQ
X-Received: by 2002:a05:690c:62c9:b0:80d:c2cc:f673 with SMTP id 00721157ae682-82022348d13mr95666737b3.20.1786030169936; Thu, 06 Aug 2026 08:29:29 -0700 (PDT)
Received: from [192.168.0.224] ([75.180.225.168]) by smtp.gmail.com with ESMTPSA id 00721157ae682-8201341ddedsm41744807b3.23.2026.08.06.08.29.29 (version=TLS1_2 cipher=ECDHE-ECDSA-AES128-GCM-SHA256 bits=128/128); Thu, 06 Aug 2026 08:29:29 -0700 (PDT)
Date: Thu, 06 Aug 2026 11:29:23 -0400
From: Nick Mathews <nick@lifelightlabs.com>
To: Gary Illyes <garyillyes@google.com>
Message-ID: <b3894d5c-c1b8-443b-b94e-168c19720d56@Spark>
In-Reply-To: <CADTQi=fy+-Z1s1Aj1KEmnTq274bS0NPTs=-2o71-17gL-XDa=A@mail.gmail.com>
References: <CADTQi=dSGTGJtczLTZ1ua77C8h4pvGhrVkz3MZV7SO280MJrKQ@mail.gmail.com> <b6abbcf2-76f6-4ace-add0-e9651512f5a7@Spark> <CADTQi=fy+-Z1s1Aj1KEmnTq274bS0NPTs=-2o71-17gL-XDa=A@mail.gmail.com>
X-Readdle-Message-ID: b3894d5c-c1b8-443b-b94e-168c19720d56@Spark
MIME-Version: 1.0
Content-Type: multipart/alternative; boundary="6a74a858_46e87ccd_3fba"
Message-ID-Hash: DXGFKSL3FGYWKB2UQMEDLMZWH5NZ2QYX
X-Message-ID-Hash: DXGFKSL3FGYWKB2UQMEDLMZWH5NZ2QYX
X-MailFrom: nick@lifelightlabs.com
X-Mailman-Rule-Misses: dmarc-mitigation; no-senders; approved; emergency; loop; banned-address; member-moderation; nonmember-moderation; administrivia; implicit-dest; max-recipients; max-size; news-moderation; no-subject; digests; suspicious-header
CC: Sauron <sauron=40google.com@dmarc.ietf.org>, web-bot-auth@ietf.org, "Mirja Kuehlewind (IETF)" <ietf@kuehlewind.net>, synack@garyillyes.com
X-Mailman-Version: 3.3.9rc6
Precedence: list
Subject: [Web-bot-auth] Re: Follow-up from IETF 126: Comments on draft-illyes-webbotauth-cbcp
List-Id: Authentication of non-human users to human-oriented Web sites <web-bot-auth.ietf.org>
Archived-At: <https://mailarchive.ietf.org/arch/msg/web-bot-auth/0KzNcVFf9_8OxKC2Tph5FpNMqJo>
List-Archive: <https://mailarchive.ietf.org/arch/browse/web-bot-auth>
List-Help: <mailto:web-bot-auth-request@ietf.org?subject=help>
List-Owner: <mailto:web-bot-auth-owner@ietf.org>
List-Post: <mailto:web-bot-auth@ietf.org>
List-Subscribe: <mailto:web-bot-auth-join@ietf.org>
List-Unsubscribe: <mailto:web-bot-auth-leave@ietf.org>
Gary, Text for all three, and on the third I think the constraint is narrower than it looks. One, the scope exclusion. I would avoid excluding by label, since a site cannot see labels at request time, and I would avoid the infancy qualifier, since a best current practice outlives the sentence that dates it. Excluding by property instead, and naming the observability gap plainly: This document describes best practices for crawlers as defined in Section 1. It does not describe best practices for automated clients that retrieve resources in direct response to a specific instruction from a human user, including the class of clients commonly called AI agents or assistants. Such clients differ from crawlers in request volume, in the relationship between the client and the person it serves, and in the expectations of both the operator and the site. Best practices for those clients are out of scope for this document. Site operators should be aware that this boundary is asserted rather than verifiable. Some operators already label user-initiated fetching distinctly from crawling in their user-agent strings, but a label is a claim, and nothing in this document's mechanisms allows a site to check it. The practices in this document may therefore, in practice, be applied to clients this document does not describe. Cryptographic identification of automated clients, under development in this working group, is the mechanism by which this boundary becomes checkable. The second paragraph is the part I would argue for keeping, and I would resist softening it. Without it the exclusion reads as though the boundary is administrable by site operators today, and it is not: it becomes administrable exactly when the working group's identification mechanisms land, which also gives the informative WG reference from issue 8 a concrete job in the text rather than a courtesy mention. Two, thank you for filing the rate limiting issue. Proposed text for the first paragraph of Section 2.3, if useful: Crawler operators MUST ensure that their crawlers are equipped with back-out logic that responds to at least the following: server error responses as defined in Section 15.6 of [HTTP-SEMANTICS]; the 429 (Too Many Requests) status code defined in Section 4 of [RFC6585], honoring the Retry-After header field where present; and connection-level failures such as timeouts and resets. Crawlers SHOULD also treat a change in the character of responses as a back-out signal rather than as content. A site under load, or applying protective measures, may begin returning challenge pages, interstitials, or truncated content while continuing to respond with a 2xx status code. A crawler that continues at the same rate in this situation adds load precisely when the site is attempting to shed it. The second paragraph is the case I care most about from the receiving side. Silent degradation is more common in practice than a clean 429, and a crawler keying only on status codes cannot see it. One addition on the machine-readable question from my earlier comments: that mechanism exists in progress. draft-ietf-httpapi-ratelimit-headers, active in the HTTPAPI working group, defines RateLimit and RateLimit-Policy header fields by which a server advertises its limits. An informative reference would let this document point crawler operators at where to look for a site's stated limits once that work publishes, without any normative dependency, under the same rule as the WG reference above. Three, on not basing a best practice on future work. Your instinct about general signature based works is right, and the rules make it cleaner than it sounds. RFC 9421, HTTP Message Signatures, is Standards Track and was published in February 2024. Referencing it normatively creates no dependency on unfinished work. That is sufficient to support a statement in Sections 2.2 and 2.5 that user-agent strings and published address ranges are self-asserted from the site's perspective and cannot be treated as authentication, and that where a cryptographic mechanism is available it is preferable. The document takes no position on which profile of it prevails. The working group's own effort can then appear as an informative reference. RFC 7322 Section 4.8.6.4 states that references to Internet-Drafts may only appear as informative references, which is exactly the tool for this: a sentence noting that work to standardize cryptographic identification for automated clients is underway in this working group, cited informatively, answers Thibault's issue 8 without making conformance depend on the outcome. Happy to open any of these as pull requests against the repository if that is easier than the list. Nick Mathews AVA Pay, Agentic Verification Architecture https://github.com/AVA-PAY/ava-pay On Aug 5, 2026 at 11:27 AM -0400, Gary Illyes <garyillyes@google.com>, wrote: > Thanks Nick! > > 1. Do you have language suggestion how to specifically exclude agents? Could we just say exactly that, "AI agents are specifically excluded"? Perhaps qualify that sentence by mentioning they are still in their infancy? > > 2. That's a good point; I missed that. Filed a github issue: https://github.com/garyillyes/cbcp/issues/19 > > 3. Thibault pointed out a while ago that we should have a reference back to the WG: https://github.com/garyillyes/cbcp/issues/8. One challenge, I think, is that we cannot base a best practice on future work, which is what this WG is doing. I'm not sure how to get around that besides referencing general sig based works. > > On Tue, Jul 28, 2026 at 6:36 PM Nick Mathews <nick@lifelightlabs.com> wrote: > > Hi Gary, > > > > Comments on CBCP from the site side, in the three areas you asked about. I run AVA Pay, an open source merchant-side verifier for automated commerce traffic (Web Bot Auth, Visa TAP, AP2). We sit where these practices get enforced, so these observations come from operating the receiving end rather than from running a crawler. > > 1. Scope: user-initiated agents will be governed by this document whether or not they are in it. > > Section 1 defines a crawler as retrieving resources "without direct human initiation of individual requests." That cleanly excludes the fastest-growing class of automated traffic we see: an agent dispatched by one person, for one task, at one site, thirty seconds ago. The exclusion is correct as a definition and unhelpful as a deployment reality, because site operators cannot tell the two apart at request time. What actually happens is that crawler policy gets applied to a shopping agent carrying a customer's intent, and a sale is refused. > > Suggestion: say explicitly whether user-initiated agents are in scope, and if not, name the venue that covers them. A sentence acknowledging that sites cannot distinguish the classes without additional signals would also motivate the rest of the WG's work nicely. > > 2. Rate control: back-out logic keyed on 5xx misses the signal that rate limiters actually emit. > > Section 2.3 requires back-out logic relying on "at least the standard signals defined by Section 15.6 of [HTTP-SEMANTICS]," which is the 5xx server error range. In practice the response to excessive automated traffic is usually 429 with Retry-After, not a 5xx, and often not even that: the more common failure is silent degradation, where a site's protection layer starts serving partial or challenge content while still returning 200. > > We hit exactly this last week. Repeated automated fetches of one domain tripped that site's rate limiter; the visible symptom was degraded content well before any hard failure. A crawler keying only on 5xx would have kept going. > > Suggestions: reference 429 and Retry-After explicitly as normative back-off signals; add guidance that a sudden change in response shape (challenge pages, truncated content, redirect to interstitials) should be treated as a back-off signal rather than as content; and consider whether a machine readable way for a site to state its own limit is in scope, since today a well behaved crawler is left to guess and a badly behaved one pays no price. > > 3. Identification: the document requires identity that cannot be verified, in the working group that is building verifiable identity. > > Sections 2.2 and 2.5 rest on user-agent strings and published IP ranges. Both are self-asserted from the site's perspective. Any client can claim to be any crawler in a User-Agent header, and IP-range publication is operationally fragile in the presence of CDNs, cloud egress, and IPv6 churn. This is the concrete pain merchants describe: roughly half of commerce traffic is now automated, most sites have no policy beyond monitoring, and the identification they are offered is a string anyone can type. > > Suggestion: keep 2.2 and 2.5 as baseline hygiene, but state plainly that user-agent and IP-range identification MUST NOT be treated as authentication, and point normatively to the signature-based work in this working group as the preferred mechanism where available. Same for the central registry idea in Section 1: a registry of names without proof-of-key reproduces the problem it is meant to solve. Anchoring it in the key-directory work being drafted here seems strictly better than a parallel list. > > Happy to contribute text for any of this, or to open the issues on GitHub if you prefer them there. > > > > Nick Mathews > > AVA Pay, Agentic Verification Architecture LLC > > https://github.com/AVA-PAY/ava-pay > > On Jul 27, 2026 at 8:42 AM -0400, Sauron <sauron=40google.com@dmarc.ietf.org>, wrote: > > > Hi all, > > > Thanks for the engagement and support during the working group session at IETF126. > > > As mentioned during the meeting, we are looking for even more feedback on the Crawler Best Practices draft: https://datatracker.ietf.org/doc/draft-illyes-webbotauth-cbcp/ / https://github.com/garyillyes/cbcp > > > Specifically, input on the scope, rate control conventions, and handling of non-malicious crawlers would be really useful to help move the draft forward. Please send any comments to the list or open issues directly on GitHub. > > > Thanks, > > > Gary > > > PS: if you have ideas where to send the following two drafts that CBCP depends on, that would be hugely welcome > > > 1. JAFAR: https://datatracker.ietf.org/doc/draft-illyes-webbotauth-jafar/ > > > 2. REP-ext https://datatracker.ietf.org/doc/draft-illyes-repext/ > > > _______________________________________________ > > > Web-bot-auth mailing list -- web-bot-auth@ietf.org > > > To unsubscribe send an email to web-bot-auth-leave@ietf.org > > _______________________________________________ > > Web-bot-auth mailing list -- web-bot-auth@ietf.org > > To unsubscribe send an email to web-bot-auth-leave@ietf.org