[SCITT] Re: Proposal: publish what the last week actually produced, as a citable interop report with all of us on it

Tiago Pinto <tiago@donttrustverify.pt> Tue, 25 August 2026 09:34 UTC

Return-Path: <tiago@donttrustverify.pt>
X-Original-To: scitt@mail2.ietf.org
Delivered-To: scitt@mail2.ietf.org
Received: from localhost (localhost [127.0.0.1]) by mail2.ietf.org (Postfix) with ESMTP id E796512EF9598 for <scitt@mail2.ietf.org>; Tue, 25 Aug 2026 02:34:42 -0700 (PDT)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=ietf.org; s=ietf1; t=1787650482; bh=GD0Tk5DeYiY3jXLNyIGD3oSFCJ0sulpqDJkijem0Eao=; h=Date:To:From:Cc:Subject:In-Reply-To:References; b=NF/pWx0qb1YJLmPlhhGylFKt8BxSdjw0jSC2DV5Gub5nXTmXcLDmbvTF/OBWS4eGR wvoYaIOMTRszBUcfpPhCLWIDhoB5g6Zdg9AQVFrEjM/SwppkXv0DRdkGu/SVCnQLv/ TtCRFfWZAP67tUqSkcWlkmHKBNJy8dciW6bKgt74=
X-Virus-Scanned: amavisd-new at ietf.org
X-Spam-Flag: NO
X-Spam-Score: -2.797
X-Spam-Level:
X-Spam-Status: No, score=-2.797 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, RCVD_IN_DNSWL_LOW=-0.7, RCVD_IN_MSPIKE_H5=0.001, RCVD_IN_MSPIKE_WL=0.001, RCVD_IN_VALIDITY_RPBL_BLOCKED=0.001, RCVD_IN_VALIDITY_SAFE_BLOCKED=0.001, SPF_PASS=-0.001] autolearn=ham autolearn_force=no
Authentication-Results: mail2.ietf.org (amavisd-new); dkim=pass (2048-bit key) header.d=donttrustverify.pt
Received: from mail2.ietf.org ([166.84.6.31]) by localhost (mail2.ietf.org [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id bQV7KF_bSfUZ for <scitt@mail2.ietf.org>; Tue, 25 Aug 2026 02:34:39 -0700 (PDT)
Received: from mail-4318.protonmail.ch (mail-4318.protonmail.ch [185.70.43.18]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature ECDSA (P-256) server-digest SHA256) (No client certificate requested) by mail2.ietf.org (Postfix) with ESMTPS id E635512EF9581 for <scitt@ietf.org>; Tue, 25 Aug 2026 02:34:38 -0700 (PDT)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=donttrustverify.pt; s=protonmail; t=1787650470; x=1787909670; bh=R30IPKqcvXxxCvJDa+i146EkwbM53Tzdi+sWfHVLE2k=; h=Date:To:From:Cc:Subject:Message-ID:In-Reply-To:References: Feedback-ID:From:To:Cc:Date:Subject:Reply-To:Feedback-ID: Message-ID:BIMI-Selector; b=TicQlmk3qLnoMeZpSHqyi7Br8Ri4vc7h16i19GcbKi3JpfyD7fFBPPqgJpRrZfC+J DG8DmxnpCfxEmjO9zaNJGZv4RnIZdGyQRaCu3TZWJ2svtvkZVg9D8GQV/oQ10SIt5K Hhx0miYFAnvk2SR6F6evpgqESKiRqEGyNkMpTNqgQNDXbngNZTB9MDeaxbfgpIBXX8 o4F8/9iWgCRMI8XuTrgLEvXhKA67+1xs7bICQTrfumIz6FypyqDEOUUyWuVcU3emCW VHyaha06S4S3aUMHY/OjdoXLZRKPe036VHXSzj+/fl5ZXlG1m8dH3zUX9k7VNDtbih EdqSRnkgIAMng==
Date: Tue, 25 Aug 2026 09:34:27 +0000
To: Joel Hillier <jhillier=40certisyn.com@dmarc.ietf.org>
From: Tiago Pinto <tiago@donttrustverify.pt>
Message-ID: <C1FtZkTY7C4LoBQhGXVQi63E1lUDt1akNAlTAC5MbvycgrFZVHw7zalPA1FoBG52fspigL3jtnNePTbqNZ5MsQQjTQRsliLpJTrZEmhUep0=@donttrustverify.pt>
In-Reply-To: <BYAPR19MB2806467E4B1E98518434FBD5ADA22@BYAPR19MB2806.namprd19.prod.outlook.com>
References: <BYAPR19MB2806467E4B1E98518434FBD5ADA22@BYAPR19MB2806.namprd19.prod.outlook.com>
Feedback-ID: 206570780:user:proton
X-Pm-Message-ID: 8386e25c613d3343deffd09930e61553bcb59e44
MIME-Version: 1.0
Content-Type: text/plain; charset="utf-8"
Content-Transfer-Encoding: quoted-printable
Message-ID-Hash: YOUQNOEBBB63MR73ESMR7WGOHLIIEWVO
X-Message-ID-Hash: YOUQNOEBBB63MR73ESMR7WGOHLIIEWVO
X-MailFrom: tiago@donttrustverify.pt
X-Mailman-Rule-Misses: dmarc-mitigation; no-senders; approved; emergency; loop; banned-address; member-moderation; nonmember-moderation; administrivia; implicit-dest; max-recipients; max-size; news-moderation; no-subject; digests; suspicious-header
CC: "scitt@ietf.org" <scitt@ietf.org>, "wdhawkins46@gmail.com" <wdhawkins46@gmail.com>, "hello@vaara.io" <hello@vaara.io>, "team@emiliaprotocol.ai" <team@emiliaprotocol.ai>, "playplay2736@gmail.com" <playplay2736@gmail.com>, "nenadvasic@protonmail.com" <nenadvasic@protonmail.com>, "Todd.Gibson@t-mobile.com" <Todd.Gibson@t-mobile.com>, "anton.sokolov@tyche.institute" <anton.sokolov@tyche.institute>
X-Mailman-Version: 3.3.9rc6
Precedence: list
Subject: [SCITT] Re: Proposal: publish what the last week actually produced, as a citable interop report with all of us on it
List-Id: "Supply Chain Integrity, Transparency, and Trust" <scitt.ietf.org>
Archived-At: <https://mailarchive.ietf.org/arch/msg/scitt/aLlCnsdmeF7-kHGqhV2Vb_L5EHo>
List-Archive: <https://mailarchive.ietf.org/arch/browse/scitt>
List-Help: <mailto:scitt-request@ietf.org?subject=help>
List-Owner: <mailto:scitt-owner@ietf.org>
List-Post: <mailto:scitt@ietf.org>
List-Subscribe: <mailto:scitt-join@ietf.org>
List-Unsubscribe: <mailto:scitt-leave@ietf.org>

Joel,

Yes, with changes.

Since you asked us to correct our own part, I have two corrections to
mine.

First, the sentence "three more contributed findings without running
anything" no longer describes my contribution accurately. On 24 August I
reproduced the CAP-1 class digest from the public tree at `0980d32` and
reported the result to the list; Pablo's message later that evening
independently grouped that work with the contributions that involved
execution. That reproduction is separate from the review I had frozen
and timestamped earlier. The first row is a review; the second is a
reproduction.

Second, that reproduction does not fit any of the three run kinds you
proposed.

I did not reproduce an author-supplied checker because I ran none of
your verifier code. I did not write an independent implementation
because I implemented no verifier. What I reproduced was a construction
specified in prose, in order to test whether the description was
sufficient to derive the published value independently.

That distinction mattered. The path root was not fixed by the
construction as stated, and Emek and I initially obtained different
values from otherwise consistent procedures. Once the root was
identified, the class digest reproduced.

If the report has a `run_kind` column, I think this deserves a fourth
category, something like `construction reproduction`. I suspect it will
be useful for more than my row.

I would also change one sentence in the framing.

I do not think "not one of them was found by reading" survives the
record. The paragraph immediately above already says that one defect was
reached analytically by one person and empirically by another. The
`withheld` contradiction was found analytically in my review, frozen and
timestamped before I had read the other responses, and Walter then
confirmed it against PV-04.

The more interesting result is not that execution beat reading. It is
that the two methods crossed. Some defects were found analytically, some
through execution, and several were independently confirmed by a
different person using the other method. I think that is both more
accurate and a stronger result.

I have two structural requests for the report.

First, I think it should be explicit about what each author is actually
vouching for. I can stand behind my own rows and the methodology I have
personally checked. I cannot personally attest to your two hundred
signature calls, Nicholas's corpus run, Henri's deposit, or any other
experiment I did not reproduce. Those should remain attributable to the
person who performed them rather than eight author names appearing to
certify every measurement collectively.

Second, I would keep the roles separate at finding level. Where
applicable, record who originated the finding, who implemented the
relevant test or repair, who reproduced it independently, and who
adjudicated it. Last week's work contains several cases where those were
different people. Flattening that into a single attribution would remove
some of the provenance that makes the result useful.

I would also widen the title slightly. Several rows are explicitly not
cross-implementation interoperability results, and people have already
been careful to say so about their own work. Something along the lines
of:

`Adversarial Implementation, Reproducibility and Interoperability Report`

would let the report describe those rows accurately without having to
qualify the title each time.

For the deposited artefact, I would freeze and hash the final text
before deposit, publish that digest to the list, and include the digest
in the record around the deposit. A reproducibility report should make
the identity of its own final bytes unambiguous.

Finally, I would keep the charter argument outside the deposited report.
The report should describe what was measured and what was observed.
Anyone who wants to use those measurements in an argument about future
charter scope can make that argument separately and in their own name.

For the author list, please use:

Tiago Pinto
Independent researcher
donttrustverify.pt

Best,

Tiago



Em terça-feira, 25 de agosto de 2026 às 03:32, Joel Hillier <jhillier=40certisyn.com@dmarc.ietf.org> escreveu:

> Hi all,
> 
> Iman has put an open invitation on the list for a compatible implementation to run one COSE test set across two implementations before IETF 127, so the working group gets a concrete result rather than an assertion. That is the right instinct and it deserves more than one pair. Something happened on this list over the last week that I don't think any of us planned, and it's worth writing down properly before it scrolls away.
> 
> What actually happened, stated at the precision I can support.
> 
> Five of us ran code somebody else wrote, at pinned commits or package digests, on different operating systems and different language runtimes, and published the results so anyone can check them. Three more contributed findings without running anything, including one review frozen and OpenTimestamps-stamped before its author had read anyone else's. Along the way:
> 
> -   Two verifiers written independently from one text disagreed on 15 of 15 documents, and the root cause was a member name the text never gave.
> -   A published schema was found to make a normative MUST unsatisfiable at two of three sites where other rules rely on it.
> -   An anchored sequence was shown to be poisonable by anyone willing to pay for one transaction, and its author's own anchor was then refused by the rule that fixed it.
> -   A conformance protocol was shown to be unsatisfiable where its rules entail one another, analytically by one person and empirically by another, inside a vector its author had written.
> -   A single exit code was found to name five different conditions, and the guard over the set of codes was blind to it by construction.
> -   A valid JSON file was found to be reported as invalid JSON, which is a false message rather than a wrong code.
> -   A manifest runner was found to enumerate what was present on a build machine while a constant at the top of it said tracked, so a file in no commit would have been hashed into a published archive digest.
> -   A COSE_Sign1 envelope digest was found to be standing for two things a relying party has to tell apart: the registration entry, and the authorization claim inside it. That is Iman's, it's published as ten source-locked cases, and it is the same shape as the exit code and the verdict vocabulary above, in a different layer of the stack.
> 
> Three of those are already repaired, and that matters more than the list. The exit-code and JSON-message findings are fixed in `@conarium-ai/core` 0.2.44 through 0.2.46, so exit 20 now names one condition rather than five and the JSONL message says what is actually true. The manifest runner is repaired in my tree, carries two mutants, and the tree it describes now verifies 30 of 30 with every enumerated file tracked by git. Anyone reproducing any of the originals needs the pinned version it was found at, which is precisely why the report has to record what was run against what rather than only what was found. A finding without its version is an anecdote a month later.
> 
> Not one of them was found by reading.
> 
> And the last one generalises, which is the argument for doing this properly. I measured signature stability across the algorithms a signing estate migrates through: one key, one signing input, 200 ordinary sign calls per leg, no adversary and no transform. The Sig_structure digest is 1 of 200 on every leg. The envelope digest is 1 of 200 on Ed25519 alone, 200 of 200 on ECDSA P-256, 200 of 200 on hedged ML-DSA-65, and 200 of 200 on a parallel Ed25519 plus ML-DSA-65 composition even though the classical leg alone is 1 of 200. Hedged and deterministic ML-DSA-65 are both FIPS 204, both verify, and are selected by a flag absent from the protected header, so two conforming signers disagree about whether the envelope digest is stable at all. The probe and its run record are pinned in my conformance tree and the numbers reproduce on two machines. Detail is in the agenda thread.
> 
> What I'd like to propose. That we publish it as a joint interop report. Not a draft, not anybody's specification, and not an argument for anyone's document. A record of what was run, by whom, against what, and what it showed.
> 
> One. A table of every run: implementation, author, what it was written from, what it was run against, the pinned commit or package digest, and the result. Using the run-kind vocabulary Iman wrote for his own row and Emek then applied to his: reproduction of author-supplied checkers, independent implementation from the text, or independent implementation with independent vectors. A row that doesn't name its kind reads as the first. Henri has since shipped that field in Vaara v1.75.0, so the vocabulary exists in a form other people can file against rather than only in this thread.
> 
> Two. The disagreement matrix. Where two implementations differed, what they differed about, and which of them the text could actually settle. The disagreements are the load-bearing part. A report recording only agreement would be worthless and everyone here knows it.
> 
> It also needs a column the last two days argued for: what pairs of implementations are not comparable, and why. Emek and Iman established on the list that two runners in different problem domains produce no cross-implementation result between them, and said so rather than letting the adjacency imply one. A report that records only the comparisons that worked would misrepresent the exercise as neatly as one that records only agreement.
> 
> Three. The findings, each attributed to whoever found it, with the version it was found at and the version it was repaired in where it has been. Including the ones people found in their own work. There are a lot of those and they are the reason the exercise is credible at all.
> 
> Four. A DOI, so it can be cited rather than linked. Henri has already deposited Vaara this way at 10.5281/zenodo.22029444, and it turns a repository URL into a reference. I'll do the deposit and the editing unless someone would rather.
> 
> Who's on it. Everyone who ran something or contributed a finding, as an author rather than an acknowledgement. On the work so far that's Walter Hawkins, Emek Can Doğru, Tiago Pinto, Henri Sirkkavaara, Iman Schrock, Pablo, Nenad Vasic and me. Anton, your ECDSA work is upstream of a repair in my document and I'd put you on it too if you want to be. Todd, Walter named the receiver vantage as yours and it belongs in the report, but I'm not putting your name on something you haven't written a word of, so tell me either way. Iman, the consent you gave Anton covers naming EMILIA in the 127 discussion; an authored row here is a different ask, so tell me rather than have me assume it carries over.
> 
> If I've mis-stated what anyone did, say so and it gets corrected before anything is deposited. I'll write the first draft and circulate it here. Nobody's name goes on it without them reading the text, and any of you can strike any sentence about your own work without giving a reason.
> 
> Why it's worth the effort. Two reasons, and I'd rather state the second plainly than have it inferred.
> 
> The first is that this list is about to have a conversation about what happens next, and the chairs have opened it: the IETF 127 agenda call went out with implementation and interoperability experience named as a category. The chartered work is close to done. RFC 9943 published in June, SCRAPI in the RFC Editor queue, and the CCF profile in Last Call. Meanwhile a substantial body of individual work has accumulated around this group and almost none of it is in the current charter. Off my own reference lists rather than a count I'd have to defend, that includes attested agent payment, disclosure evidence, canonical action identifiers, payload binding, agent accountability composition and conformance, agent records, web bot auth in three parts, receipt formats, and two of mine. Three more landed in nine days, two of them within 48 hours of my own last revision. When somebody asks whether there is work here to charter, the honest answer is either an opinion or a measurement, and right now we are the only people who can hand over a measurement.
> 
> The second is that it is good for every one of us commercially, and pretending otherwise would be silly. I sell verification. Several of you sell adjacent things. A public, reproducible, adversarial interop study with eight names on it is worth more to all of us than any of us saying the same thing alone, and it is worth more precisely because none of us controls it.
> 
> What it is not. It is not a vehicle for ARP or for CAP-1, and I'll take both out of the framing entirely if anyone thinks they're leaning on it. My drafts appear in it exactly as everyone else's do: as things that were run, with what broke.
> 
> Say yes, no, or yes-with-changes. If the answer is no I'll drop it without argument, and the work stands on its own in the archive either way.
> 
> Joel