[SCITT] Re: Deterministic encoding fixes the value, not the type: a four-part check for any digest a second party has to recompute

e.dogru@conarium.dev Fri, 21 August 2026 11:39 UTC

Return-Path: <e.dogru@conarium.dev>
X-Original-To: scitt@mail2.ietf.org
Delivered-To: scitt@mail2.ietf.org
Received: from localhost (localhost [127.0.0.1]) by mail2.ietf.org (Postfix) with ESMTP id 53FAA12D4D27E for <scitt@mail2.ietf.org>; Fri, 21 Aug 2026 04:39:46 -0700 (PDT)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=ietf.org; s=ietf1; t=1787312386; bh=LQqtUeareCGt//FI98IOgrZ9qqLy9OgQE4YJnSjthOs=; h=Cc:From:In-Reply-To:References:Subject:To:Date; b=DlnsvPPQGKoC8GObe/5TrT2qduloMg3rz7GVJIBj+XWAVNyvfIRp+kbtcXy9CEh+4 jdVswMmogFnsm811ALP4JslaAKronhKN6RbDtq5AaId6T5s2tWbm8miqED7Va0ieno 0rjmV/z5pH8zDHKtgdhx4E/iZHML17ZUv+cob3QI=
X-Virus-Scanned: amavisd-new at ietf.org
X-Spam-Flag: NO
X-Spam-Score: -1.618
X-Spam-Level:
X-Spam-Status: No, score=-1.618 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, HTML_MESSAGE=0.001, HTML_MIME_NO_HTML_TAG=0.377, MIME_HTML_ONLY=0.1, RCVD_IN_DNSWL_NONE=-0.0001, RCVD_IN_MSPIKE_H4=0.001, RCVD_IN_MSPIKE_WL=0.001, RCVD_IN_VALIDITY_RPBL_BLOCKED=0.001, RCVD_IN_VALIDITY_SAFE_BLOCKED=0.001, SPF_HELO_NONE=0.001, SPF_PASS=-0.001] autolearn=no autolearn_force=no
Authentication-Results: mail2.ietf.org (amavisd-new); dkim=pass (2048-bit key) header.d=conarium.dev
Received: from mail2.ietf.org ([166.84.6.31]) by localhost (mail2.ietf.org [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id G1Ae2kw0ihZ8 for <scitt@mail2.ietf.org>; Fri, 21 Aug 2026 04:39:45 -0700 (PDT)
Received: from barn.pear.relay.mailchannels.net (barn.pear.relay.mailchannels.net [23.83.216.11]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature ECDSA (P-256) server-digest SHA256) (No client certificate requested) by mail2.ietf.org (Postfix) with ESMTPS id 94F0E12D4D278 for <scitt@ietf.org>; Fri, 21 Aug 2026 04:39:45 -0700 (PDT)
X-Sender-Id: hostingeremail|x-authuser|e.dogru@conarium.dev
Received: from relay.mailchannels.net (localhost [127.0.0.1]) by relay.mailchannels.net (Postfix) with ESMTP id 7EFEA4013EA for <scitt@ietf.org>; Fri, 21 Aug 2026 11:39:39 +0000 (UTC)
Received: from fr-int-smtpout30.hostinger.io (trex-green-7.trex.outbound.svc.cluster.local [100.96.17.28]) (Authenticated sender: hostingeremail) by relay.mailchannels.net (Postfix) with ESMTPA id E86854020E7 for <scitt@ietf.org>; Fri, 21 Aug 2026 11:39:38 +0000 (UTC)
X-Sender-Id: hostingeremail|x-authuser|e.dogru@conarium.dev
X-MC-Relay: Good
X-MailChannels-SenderId: hostingeremail|x-authuser|e.dogru@conarium.dev
X-MailChannels-Auth-Id: hostingeremail
X-Wipe-Skirt: 2e0ce4b5296d2114_1787312379459_2698705638
X-MC-Loop-Signature: 1787312379459:2655127346
X-MC-Ingress-Time: 1787312379459
Received: from fr-int-smtpout30.hostinger.io (fr-int-smtpout30.hostinger.io [148.222.54.7]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384) by 100.96.17.28 (trex/8.0.2); Fri, 21 Aug 2026 11:39:39 +0000
Received: from localhost (34.86.89.34.bc.googleusercontent.com [IPv6:2a00:1d35:17a4:d900:51f5:85e2:e74c:e841]) (Authenticated sender: e.dogru@conarium.dev) by smtp.hostinger.com (smtp.hostinger.com) with ESMTPSA id 4hRJHs1G9Cz2xNp for <scitt@ietf.org>; Fri, 21 Aug 2026 11:39:37 +0000 (UTC)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=conarium.dev; s=hostingermail-a; t=1787312377; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=SOffz9eO77dS3I7ACwzHe81NLXNNHdcaq+iFUj3hIoU=; b=ev6HNuGRto08IshwQroGFmj7eWCDqY5M3Ba0129PljVh7rmS0R2LnJYQZrV9/mwPJK78yw rXyEZ6aybEcz/ride+ZY/a0JQwjTNEgYWKgv3kf2I76kKrODPkTaP7L7rKF1IguOREF6Y+ W76ZeOyH0hoe+kGeSv5pzQ+SsqmgqWi4mBvD+26Vw+BRcUgV+oKabZRGc49xtZrZiqmXOF b7DWFLppOdK+nf5dirbu2gMG8O8JvIK0X0IZiABtqA8iYkBOa90zLqDgDNSVSnkIkaIFzz 435ShRRo6EOW075yRHUjafiA5Te0HWHrKerLmg8H0GAa9W2fAZzeowbpgthQeQ==
Content-Transfer-Encoding: quoted-printable
Content-Type: text/html; charset="utf-8"
From: e.dogru@conarium.dev
In-Reply-To: <BYAPR19MB280695565F7402EBF09819BCADA42@BYAPR19MB2806.namprd19.prod.outlook.com>
Message-Id: <1787312368130682019.1787312368@conarium.dev>
Mime-Version: 1.0
References: <BYAPR19MB28065B9D5E107213DDE3163DADA42@BYAPR19MB2806.namprd19.prod.outlook.com> <giOlS6Y51IJfg3rjyjyxlQa-YeZkeTpNKGqjMvVlFmY5-chm7M3h36AkSj29RzEdCIiOkMWcucrsZ2PnLCA0upE9ek2Ppg9k2m-xgtaQzNw=@vaara.io> <BYAPR19MB280695565F7402EBF09819BCADA42@BYAPR19MB2806.namprd19.prod.outlook.com>
To: jhillier=40certisyn.com@dmarc.ietf.org
Date: Fri, 21 Aug 2026 11:39:37 +0000
X-CM-Analysis: v=2.4 cv=V5Av0vni c=1 sm=1 tr=0 ts=6a8838f9 a=RddCdUNZqxBAE8jYSUBS9Q==:617 a=9Rc7G7dILARjqDwS:21 a=xqWC_Br6kY4A:10 a=IkcTkHD0fZMA:10 a=48vgC7mUAAAA:8 a=lebMqex1blHY0lBoM0oA:9 a=nmBwlCPhS-vbLhUa:21 a=frz4AuCg-hUA:10 a=QEXdDO2ut3YA:10
X-CM-Envelope: MS4xfKWQbsQNmDPEhlRNtbtV4fGJGWSv4d9fkGrvhUOMV0j1wJ7M+EZH+UL+nQZuzgbRA0Bp9hiFDOTyOw0K+a80Bb6l2xfaTPU5u8mjk00KtgW7M6aN7MCJ sPowGzaUsiJfgPHMoopWRgo4FqIN3xRsZuRpN54jnnnTRm7L/MRfPB8YnDYeKFP110H40WLGq3aT+recbvUdJzrIIf+IuAiynWEFHflBfz+7srEpv6QkEvxO 5FG5L23pLAmJfq70ugBylA==
X-AuthUser: e.dogru@conarium.dev
Message-ID-Hash: AJEXM3PEX3PP6FTT2T2GRUB26JRU7EUJ
X-Message-ID-Hash: AJEXM3PEX3PP6FTT2T2GRUB26JRU7EUJ
X-MailFrom: e.dogru@conarium.dev
X-Mailman-Rule-Misses: dmarc-mitigation; no-senders; approved; emergency; loop; banned-address; member-moderation; nonmember-moderation; administrivia; implicit-dest; max-recipients; max-size; news-moderation; no-subject; digests; suspicious-header
CC: scitt@ietf.org, hello@vaara.io, wdhawkins46@gmail.com, anton.sokolov@tyche.institute, tiago@donttrustverify.pt, Todd.Gibson@t-mobile.com
X-Mailman-Version: 3.3.9rc6
Precedence: list
Subject: [SCITT] Re: Deterministic encoding fixes the value, not the type: a four-part check for any digest a second party has to recompute
List-Id: "Supply Chain Integrity, Transparency, and Trust" <scitt.ietf.org>
Archived-At: <https://mailarchive.ietf.org/arch/msg/scitt/Aarcoxo7WrOT4a58E0jy7CLDIgI>
List-Archive: <https://mailarchive.ietf.org/arch/browse/scitt>
List-Help: <mailto:scitt-request@ietf.org?subject=help>
List-Owner: <mailto:scitt-owner@ietf.org>
List-Post: <mailto:scitt@ietf.org>
List-Subscribe: <mailto:scitt-join@ietf.org>
List-Unsubscribe: <mailto:scitt-leave@ietf.org>

Joel, Henri,

*A runner must exercise the class, not an enumeration of previously observed
defects.* I want to report against that rather than agree with it, because we
published a failure of exactly that standard today, against our own vectors.

I should be exact about the order, because it would be easy to imply something
better than what happened. The section was written from an independent
implementer's report before I had read this thread, and it shipped several hours
after your message. Not foresight — two readings of the same problem arriving on
the same day.

Our conformance set is thirteen frozen cases, and it publishes the canonical
hashes so a second implementation can compare bytes without our key. Every case
is ASCII-keyed, with integer and string values. A naive sorted-key serialiser
reproduces all thirteen hashes. So passing our vectors does **not** demonstrate a
conforming JCS implementation, and until today nothing said so — a reader who
passed thirteen out of thirteen would have concluded a property that held over a
sample of one.

The spec states it now: RFC 8785 §3.2.2.3 number formatting and the UTF-16 code
unit ordering over non-ASCII keys are unexercised, and a receipt carrying a
float, a large integer or a non-ASCII key is where two implementations first
disagree. Naming it is not fixing it. Your framing says what fixing it looks like
— generate the class and assert invariance, and where the class is unbounded, say
in the transcript that you sampled — and that is the shape the next vectors have
to take, rather than more members of the same ASCII family.

On your relocation argument, one data point from a JSON-native format that has
been running.

You are right that JSON removes the type ambiguity and moves the gap to
rendering. We met both members of it and closed them by construction rather than
by rule, which is worth reporting because it turned out cheaper than it sounds:

- the content digest is never taken on trust. It is recomputed from the canonical
  bytes and compared, so any of your five renderings is a mismatch rather than a
  second valid spelling of the same value. Where a digest is carried as a member
  instead of recomputed, the schema pins its rendering explicitly — which is your
  rule, arrived at one member at a time rather than stated once;
- the signature is base64 and a non-canonical variant is **refused rather than
  normalised**: the decoder re-encodes what it parsed and rejects anything that
  does not round-trip to the original string.

Neither was chosen for this argument. The second exists because normalising on
ingest is how you manufacture a twin, which is the failure Anton and Henri are
looking at in the ECDSA thread from the other end.

The one we have not closed is your number/string boundary. Nothing in our format
pins whether a count is `25` or `"25"` except the schema, and the schema is not
in the preimage. That is a real hole and I would rather record it here than find
it later.

Emek


On Fri, Aug 21, 2026 at 12:06 AM Joel Hillier <jhillier=40certisyn.com@dmarc.ietf.org> wrote:

Hi Henri,

Your per-digest framing is better than mine and I'd adopt it as the requirement. "A specification that cites 4.2.1 once at the top has told a reader nothing about any particular preimage" is the sentence I should have written. Per document, the check is a posture.Per digest, it's an obligation with a visible hole when it isn't met.

Working through it on my own text overnight, I think it goes one step further than either of us put it, and the step is worth taking because it changes what a conformance runner is for.

The four gaps aren't four things. They're one thing in eight forms.

Signature encoding. CBOR type choice. Collection order. Absence, omitted or nulled. Text normalisation. Reference normalisation. Binary rendered into a text encoding. Framing, a concatenation against an array against a map.

In every one of those, two byte strings carry the same information and a digest over them differs. The information is invariant and the representation isn't, and the digest got defined over the representation. That's the whole family, and once you see it thatway the requirement stops being a checklist and becomes a property: a digest must depend only on what its input means, and not on any choice of encoding that carries no meaning.

Physics has a name for a quantity that must not depend on a representational choice carrying no information, and more usefully it has a habit that goes with the name: when you find one of these, you don't patch the place it showed up. You identify the freedomand quotient it out.

Which gives the conformance obligation, and this is the part I'd argue for hardest.

A runner must exercise the class, not an enumeration of previously observed defects.

A runner that checks a digest against the two ECDSA signature encodings it has heard of confirms what its author already knew. A runner that enumerates the representation class and asserts the digest is constant across it finds the member nobody had thoughtof. Same for the other seven: not "does it handle NFD", but "generate the normalisation class and assert invariance". Not "is the set sorted", but "permute and assert".

Where the class is unbounded the runner has to sample it, and it has to say in its transcript that it sampled rather than enumerated, so a reader isn't told a property holds over a class when it holds over a sample of one. That last part is the same failureas a probe computing sha256(X) == sha256(X), one level up: a test that reports a property it never exercised.

I've put this in as a section. Eight named classes, a requirement that every new digest state which class each preimage element belongs to and which member the preimage takes, and the same obligation on any IANA registration that introduces a value carriedinto a digest. It also states what it doesn't establish, which is anything about correctness: a digest can be perfectly invariant and taken over the wrong elements.

On gap 1 and JCS, you're right, and I'd put the correction one layer down rather than accept the removal.

JSON has no tag 0, no tag 32 and no bstr, so within the value model the type ambiguity is gone by construction. That's real and it's a genuine advantage of canonicalising over JSON.

What it does is relocate the gap rather than close it, and your own fourth-gap example is the proof. JSON has no byte string, so anything binary has to be rendered as text, and JSON gives no rule for which rendering. Base64, base64url, padded, unpadded, lowercasehex: five renderings of one digest, all valid JSON strings, all surviving JCS unchanged, because JCS normalises the string and has no view on how the bytes became one. The number/string boundary is the same shape. JCS 3.2.2.3 pins how a number renders, andnothing pins whether a threshold of twenty-five is the number 25 or the string "25".

So for a JCS document I'd state it as: every member whose value isn't natively a JSON scalar carries its text rendering as part of its definition. Same class, different layer. And it lands hardest exactly where you found it, at the boundary where a CBOR-nativedigest has to be recomputed by a JSON-native implementation, which isn't an encoding preference either side can resolve alone.

Gap 3 is the one where I think you've got the better of the argument for your case, and I wouldn't generalise it.

The flag-day point is right and I hadn't weighed it. Encoding an absent optional as null makes every digest computed before the member existed incomparable the day it's introduced, and omission is what lets a member arrive without breaking the corpus behindit. For a schema that's still growing, that's decisive.

The cost is specific, and it's the same failure this list has spent three days on in the other thread. Omission makes "the party held this and declined to supply it" and "the party's software predates the field" the same bytes. A null says the position existsand was empty. An omission says nothing at all, and a member that was never created leaves no gap to find. That's closing omission, one layer down, inside a preimage.

Which suggests the requirement isn't null or omit, but this: if you omit, carry something in the preimage that tells a reader which members were defined when the digest was taken. A profile or schema version inside the preimage makes an absent member attributablerather than ambiguous, and gives you back the flag-day-free upgrade that omission bought you. ARP gets this by accident rather than design, since the Policy-Version Hash sits in the preimage and pins the rule set the answer was produced under, so an absentfield is datable. I'd rather it were deliberate, and I'd put that in shared text as the condition under which omission is safe.

The fourth-gap practice is the part I'd want other people to copy. Your vectors pin the field mapping and never the scope digest, because Vaara ingests the decoded map while the digest is over the CBOR bytes, and you publish that limit instead of a numberthat looks comparable and isn't. The general rule is worth saying plainly: a vector that pins a decoded field mapping is a vector about ingestion, not about identity. A suite that doesn't distinguish the two will report agreement between two implementationsthat would compute different digests over the same input. The limit is worth more than the number would have been.

The check has teeth on my own text, which I'd rather demonstrate than assert. Run across all fourteen of ARP's digest constructions, it found gap 2 in exactly the component you and Pablo proposed lifting into the shared statement, plus a disjunctivestate identifier with no discriminator and a value carried twice with no equality rule in a document that has the correct sentence about that and applies it to the other two headers. Details are on the omission thread, because they needed to be there beforeanyone lifted it.

It also found a sixth signature carrier the earlier sweeps had walked past, which is on the ECDSA thread, and which is the argument for stating the rule over the class rather than over an enumeration: the enumeration was right until someone added a field.

Joel