[Atp] Re: Repository and Synchronization Draft Text

Phil Schleihauf <uniphil@gmail.com> Mon, 06 July 2026 22:47 UTC

Return-Path: <uniphil@gmail.com>
X-Original-To: atp@mail2.ietf.org
Delivered-To: atp@mail2.ietf.org
Received: from localhost (localhost [127.0.0.1]) by mail2.ietf.org (Postfix) with ESMTP id 2B66C11147618 for <atp@mail2.ietf.org>; Mon, 6 Jul 2026 15:47:07 -0700 (PDT)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=ietf.org; s=ietf1; t=1783378027; bh=uqIZc9IvgoItqxn9s5kf0pKJvLknZVXVCDpzI2OQB6g=; h=References:In-Reply-To:From:Date:Subject:To:Cc; b=mQlKiV/oYO3qwfJEY68ZrPxNEIIWVfs7vs7xbj9m3HjxgBbXmzqWZkqwDINRovA+f 8UcU9WY3wwwkMLedY0k3+gIZa0QJv3SPNzBGjknj7uAQi75H/9PWNFjp4HFGX3jK/R xzNK1WKm3e5BoaFSoL2W5KVBlmbC8EM4I2KvX9CM=
X-Virus-Scanned: amavisd-new at ietf.org
X-Spam-Flag: NO
X-Spam-Score: -2.088
X-Spam-Level:
X-Spam-Status: No, score=-2.088 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, FREEMAIL_FROM=0.001, HTML_MESSAGE=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_HELO_NONE=0.001, SPF_PASS=-0.001, T_KAM_HTML_FONT_INVALID=0.01] autolearn=unavailable autolearn_force=no
Authentication-Results: mail2.ietf.org (amavisd-new); dkim=pass (2048-bit key) header.d=gmail.com
Received: from mail2.ietf.org ([166.84.6.31]) by localhost (mail2.ietf.org [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id ebvUkSJW_4ts for <atp@mail2.ietf.org>; Mon, 6 Jul 2026 15:47:05 -0700 (PDT)
Received: from mail-yx1-xb131.google.com (mail-yx1-xb131.google.com [IPv6:2607:f8b0:4864:20::b131]) (using TLSv1.3 with cipher TLS_AES_128_GCM_SHA256 (128/128 bits) key-exchange X25519 server-signature ECDSA (P-256) server-digest SHA256) (No client certificate requested) by mail2.ietf.org (Postfix) with ESMTPS id 0AA70111449DA for <atp@ietf.org>; Mon, 6 Jul 2026 15:42:52 -0700 (PDT)
Received: by mail-yx1-xb131.google.com with SMTP id 956f58d0204a3-6662551100bso5301241d50.0 for <atp@ietf.org>; Mon, 06 Jul 2026 15:42:52 -0700 (PDT)
ARC-Seal: i=1; a=rsa-sha256; t=1783377771; cv=none; d=google.com; s=arc-20260327; b=hEduoUgzWwJu0uLXRUBEuJq7fO+yxDKNR94PgloFqogKsek8mbliQCr4d0IZppK7Sx DsNTtLIfQu9bgt1AT939izYfUn1tcHYAO8TK9yzyOyISTbVcSk5Jg9mLUJh37jjZmvgW mzOe34dOa+IDq/7QEcl0TWnO41p8qtpe3R13n7QQmdCk0b5kPzp1SWIURd8UxvwZSapQ Kak0V6GpSo87lB/BLXWSc/dO+tNtwLlUdZ2V/qJ772Xlch3VTLuDjHUi4kjJ0vrfKK8N j+OTBTXtRVO2gGIWTJUDirNdFv9JklKoNzNzkcK0D0WBWzZQZ/EiqZOCYB0KD1btx4/o +fDQ==
ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=arc-20260327; h=cc:to:subject:message-id:date:from:in-reply-to:references :mime-version:dkim-signature; bh=uqIZc9IvgoItqxn9s5kf0pKJvLknZVXVCDpzI2OQB6g=; fh=lfsbB+MBZbIrr5Ww92OSkYm7YyqNTZirdMYwnw8vgtA=; b=VtYlKijQTgcExx6euMw6XFKIzZhDzvHj0hdASwPX6ILloxouBJ84AIut859ih1dtb6 ljNwYoQrpGw7nTpchUk6+bin2AUq3VhPahVHln9RA7J0LQcDlFCVYg8H1GQezFQgMYFI zXdErgjd5+hp5FfYgGcGjUZ83+O6TW60EZns6c1/9c9JN59ucJINiptYIkQtOVe0Wvxc V68fsUXapoiYrT7W+xey+oWHYYoCCxce1b914tPSQ04rraElRyMu4vZnhtOD9SuKF/Ne P7YY4hJZr+SFRnbQsaixwLA91TDVkqAjK1cwFZOsqCjEzA+jasB5x0YXuDFtUyO4nBEq PuRw==; darn=ietf.org
ARC-Authentication-Results: i=1; mx.google.com; arc=none
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1783377771; x=1783982571; darn=ietf.org; h=cc:to:subject:message-id:date:from:in-reply-to:references :mime-version:from:to:cc:subject:date:message-id:reply-to; bh=uqIZc9IvgoItqxn9s5kf0pKJvLknZVXVCDpzI2OQB6g=; b=FXKe8Hi0doxxcFwwcvRItT2a/x5RY9bJJHfxxb5cD5Pm362BjyDROUwhuKjt/xzXC3 O2UjWnwDZYIJjHYM0dgW30IR23HhofkNqdOT9U32+6NMq/MeZT7s1DoPY878ff8qPuIs mZi4hWQN6wWyAwWdZ8oYkDmR/zmj/+JCNhybYTBI5x/g6+84jkPljwzTlZCLZg/Hd+wW NBHtfDWORVrXP0e150FHLv+5v+RZpYiirk4Oi98hWKB78yAtx4SQt3EnviGQYbSjSQZ+ Cr/JTnbMMZUjx8tVq1bmim03EUnWpi8j9tYhmEhFyCBCZt6XLYZpCuP7zxIgd6Dj3hjm ZCXg==
X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1783377771; x=1783982571; h=cc:to:subject:message-id:date:from:in-reply-to:references :mime-version:x-gm-gg:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to; bh=uqIZc9IvgoItqxn9s5kf0pKJvLknZVXVCDpzI2OQB6g=; b=EymDraxerPVXtmtsiInjJ6X8kRE0U/M+F5ZkWuwhLbMLkapAqFtV1KGgSnHGvwkJFQ /U6sUqqSZlUuun0o2vf4lcNpc3M+qOAx1qtJL6rJKgv0MTLLUMzRrobCjdOmZAXK9PRk 2J/ko04r1OzYq52+M8ZSTS7it3cjR9TcJ0RZZO29Jg2kzWXomnij91QoQrPlqBpZLJZs XBm3eq4NB7rlTT4vLvUNo8vJXtTDP6lWlL0MqvjoalLK/8DRDekYsBtGMY9Td88rekuv FgPOwtBmnxeLj6he/MDqrrZIevjlmGCyUAMXnc/dTEvjx1i1bgXfVF0ywKieDtzU8gE8 vAjg==
X-Forwarded-Encrypted: i=1; AHgh+RqlUVNgMYwa8n63++7zKUDmy6qeZn32FxGPak+CUK0NDZeVPVcBZGv8tsJ1muVegnJr5vc=@ietf.org
X-Gm-Message-State: AOJu0Yw5FbaIVLtOQ7wU9cd+GZDIwqutb/ZlLQRgBTMesveJaieAiaCP SBzTYKuR2rmrwhSB9E8HcU27uvbz0bQXsVcw10t/5YRALHhHM+e2g1iQehaFSmjMwBNK2OE+gZL U8s87AnwL8t9m1hG8Hwf3jZWHmgaaFi0=
X-Gm-Gg: AfdE7cl2+S7WOVybHk3ebhX+GJhuAYYgDCGY8oIFZ95paL26Y1OqogrV3nV98yxh83y ct+UTs3w9PmbQ4cIWNxyEhhSQt5bg2gsWmIqjJetOr1lUShBmYN4BKCGIAhsmtpG95diWAF/YcA iqhULKJARp+gNvTBNxCpqMFDMbjgJnEs0hM/3foZ4AWcha4M4UY2nk26u/ZV8WvuEPmGeBjJHQk zz/0i+hrLMMgFaK+ETQTAx+o7GFbskHt+eKAKVRB4CEzdxBNFs2HFCbGfViDPSyfQpk+eSgnvew rwFP/gIRaw45DHKZyviOq1HTO5kz
X-Received: by 2002:a05:690e:150c:b0:664:ae68:ca0f with SMTP id 956f58d0204a3-6677febd69amr2072658d50.85.1783377771440; Mon, 06 Jul 2026 15:42:51 -0700 (PDT)
MIME-Version: 1.0
References: <CABFYohhdpHg4xBHz7Z15t0DbUPMYGEoPTPBvLuT8z6MK=2EbtA@mail.gmail.com> <CAP9=HBEk=u+6R5GdMWDi=_26zg3U1KwkT7M_FC_OmL+kn1oFqQ@mail.gmail.com> <CABFYohg7BXMicg596S_7DmVuz-b5Ljf-JEYzr5jApDtOmXF7vw@mail.gmail.com>
In-Reply-To: <CABFYohg7BXMicg596S_7DmVuz-b5Ljf-JEYzr5jApDtOmXF7vw@mail.gmail.com>
From: Phil Schleihauf <uniphil@gmail.com>
Date: Mon, 06 Jul 2026 18:42:15 -0400
X-Gm-Features: AVVi8CfNwdHAk4GcdNdFsRtgzvHYBrXguY18wc0iWq2jDOF5X89pFea9-4kcGxM
Message-ID: <CAP9=HBHMqxPSYujMRU6CXWz3pDVxToE3zLTgPvGs=rCUtX635A@mail.gmail.com>
To: Bryan Newbold <bryan@blueskyweb.xyz>
Content-Type: multipart/alternative; boundary="000000000000912df50655f8fef3"
Message-ID-Hash: NAG4YH56R7FKCQ6GXCNSX5QKWA2ZV7BN
X-Message-ID-Hash: NAG4YH56R7FKCQ6GXCNSX5QKWA2ZV7BN
X-MailFrom: uniphil@gmail.com
X-Mailman-Rule-Misses: dmarc-mitigation; no-senders; approved; emergency; loop; banned-address; member-moderation; nonmember-moderation; administrivia; implicit-dest; max-recipients; max-size; news-moderation; no-subject; digests; suspicious-header
CC: Bryan Newbold <bryan=40blueskyweb.xyz@dmarc.ietf.org>, atp@ietf.org
X-Mailman-Version: 3.3.9rc6
Precedence: list
Subject: [Atp] Re: Repository and Synchronization Draft Text
List-Id: Authenticated Transfer <atp.ietf.org>
Archived-At: <https://mailarchive.ietf.org/arch/msg/atp/7H0QNLEOOxjbtWP8XCeBriw-OdM>
List-Archive: <https://mailarchive.ietf.org/arch/browse/atp>
List-Help: <mailto:atp-request@ietf.org?subject=help>
List-Owner: <mailto:atp-owner@ietf.org>
List-Post: <mailto:atp@ietf.org>
List-Subscribe: <mailto:atp-join@ietf.org>
List-Unsubscribe: <mailto:atp-leave@ietf.org>

> I'm not sure I understand this part. If a consumer has buffered commit
messages for revs C and D, and it fetches a full repo and gets B, that
should work out fine (because commit C will reference B in the 'since'
field). If it got anything earlier it would fail. If it got C or later it
would work.

Yeah sorry, I meant the case where the consumer buffers C and D, wasn't
aware of rev B, receives repo data at A, and fails. (eg. Tap drops commits
when it puts a repo in "desynchronized" state, and starts buffering after
the transition to "synchronizing", so imagine it drops B before it tries to
resync).

In the simple-upstream case, I suppose your simple caching scenario holds:
even with different transports (websocket vs http) a consumer that begins
buffering before emitting its getRepo request, to a server that updates its
known repo rev visibly-to-http-endpoints before emitting it on the
firehose, won't race. I think.

But, that's for the boring case --

> And further this is also an issue if alternative mirrors are being used.
Eg, sources which are "siblings" as opposed to "upstream" in the data flow.

I hope these are first-class use-cases for atp too. We have all this nice
authentication stuff, and unreliable PDSes (and no current caching relays),
and the desirable option for multiple-upstream (multi-relay, relay + some
PDSes, ...) graphs, etc.

> I agree that the current expected behavior is that the consumer would
mark the repo with status "desynchronized" (per Synchronization 2.1) and
would try to re-synchronize again later.

(i think that should be S4.6?)

(This might be getting into implementation-level details instead of spec
concerns) if a first-buffered-commit-after-resync has a `since` parameter
*after* the just-resynced data's rev, then it probably lost the race. A
consumer could immediately retry the resync (retaining the existing
buffered commits), knowing the minimum rev it actually needs, needs for it
to succeed, to avoid restarting the whole race again later.

> a new optional query parameter to the HTTP endpoint for fetching full
repos

I really like this direction! However, I think it runs into the same
awareness-of-rev-B thing -- can a consumer know which rev is its minimum at
that point?

One way out of some of this might be for consumers to track the
latest-seen-rev of (valid, authenticated) firehose commits, while dropping
them, while that repo is in desynchronized state. It would be able to make
that request, or poll, or even just abandon a resync early if it's with
too-old-data.


On Tue, Jun 30, 2026 at 10:07 PM Bryan Newbold <bryan@blueskyweb.xyz> wrote:

> On Sat, Jun 27, 2026 at 8:24 AM Phil Schleihauf <uniphil@gmail.com> wrote:
>
>> Regarding Sync 4.6, Resynchronization,
>>
>> There is a data race between fetching the repository structure and
>> buffering commits to replay after completing synchronization. Suppose a
>> producer emits revs for a repo
>>
>> A => B => C => D
>>
>> A consumer realizes it needs to resynchronize upon observing rev B. It
>> then requests the repository data structure. There are no guarantees
>> provided about which version it will get. It buffers events for revs C + D
>> for replay.
>>
>> If it receives repository data from rev B, C, or D, it can successfully
>> process the buffered (and future) events. But, it might receive data from
>> rev A, in which case, processing the buffered event for rev C will fail.
>>
>> This is rare in practice today, because upstream producers (that i'm
>> aware of) in the main production atproto network always redirect resync
>> requests to the canonical hosts, which presumably always respond with the
>> latest rev they have. Latency introduced by relays works in our favour
>> against the race. But, the this section explicitly carves out scenarios
>> where latency in the other direction should be expected:
>>
>> > Direct upstreams MAY coalesce and cache snapshot requests, or redirect
>> consumers to alternative sources
>>
>> A cached snapshot or public mirror service could return a snapshot at rev
>> A when a consumer desynchronizes at rev B, causing the resync to fail.
>>
>
> I think that this would not be an issue in the case of simple caching. If
> a producer (eg, a relay) is retaining a cache of the full repository CAR
> file for an account, it would presumably invalidate that cache as soon as
> it receives notice of a new commit revision from upstream. As long as the
> cache is invalidated before the commit (or sync) message is re-emitted to
> downstream consumers, I think there is no problem.
>
> I think you are correct in the case of coalescing: in that scenario the
> timing is spread out and the result of an earlier upstream request could be
> returned to a later downstream consumer.
>
> And further this is also an issue if alternative mirrors are being used.
> Eg, sources which are "siblings" as opposed to "upstream" in the data flow.
>
>
>> Further, in this example, the consumer *knows* that it needs a snapshot
>> for at least Rev B for the resync to succeed, but that's not necessarily
>> the case. For a "scheduled resync" (eg., if the first one failed for any
>> reason, and the consumer applies some arbitrary backoff delay), the
>> consumer might not know which "minimum rev" it needs in order to succeed on
>> the next rev following resync.
>>
>> ie., it begins resync, buffering revs C and D, but has no idea about the
>> existence of rev B.
>>
>
> I'm not sure I understand this part. If a consumer has buffered commit
> messages for revs C and D, and it fetches a full repo and gets B, that
> should work out fine (because commit C will reference B in the 'since'
> field). If it got anything earlier it would fail. If it got C or later it
> would work.
>
>
>> For all of this, the consumer's handling probably boils down to "just
>> retry". While the aggregate rate of events on the network may be high, the
>> per-repo interval is expected to be low enough to eventually succeed, even
>> in the presence of slightly stale data from the resync process. Maybe some
>> of this is worth being explicit about in the spec?
>>
>
> I agree that the current expected behavior is that the consumer would mark
> the repo with status "desynchronized" (per Synchronization 2.1) and would
> try to re-synchronize again later.
>
> It seems like some of the issues you are raising could be addressed by
> adding a new optional query parameter to the HTTP endpoint for fetching
> full repos (aka, /xrpc/com.atproto.sync.getRepo, though the current draft
> text doesn't name that endpoint). If that endpoint took a query param
> 'minimumRevision' (just spitballing), the upstream could return an error if
> it didn't have access to a full CAR with that version or higher. Or, the
> consumer could call the 'com.atproto.sync.getRepoStatus' endpoint
> repeatedly until the returned 'rev' is greater or equal to what it is
> expecting (this would be cheaper than fetching the full CAR).
>
> I think this probably warrants further discussion.
>
> --bryan
>