[Cellar] Re: AD Evaluation: draft-ietf-cellar-tags-18
Orie <orie@or13.io> Mon, 07 July 2025 16:28 UTC
Return-Path: <orie@or13.io>
X-Original-To: cellar@mail2.ietf.org
Delivered-To: cellar@mail2.ietf.org
Received: from localhost (localhost [127.0.0.1]) by mail2.ietf.org (Postfix) with ESMTP id 445E4401FE65 for <cellar@mail2.ietf.org>; Mon, 7 Jul 2025 09:28:38 -0700 (PDT)
X-Virus-Scanned: amavisd-new at ietf.org
X-Spam-Flag: NO
X-Spam-Score: -2.099
X-Spam-Level:
X-Spam-Status: No, score=-2.099 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, HTML_MESSAGE=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_HELO_NONE=0.001, SPF_PASS=-0.001] autolearn=ham autolearn_force=no
Authentication-Results: mail2.ietf.org (amavisd-new); dkim=pass (2048-bit key) header.d=or13.io
Received: from mail2.ietf.org ([166.84.6.31]) by localhost (mail2.ietf.org [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id 2nlpsSSA-hkU for <cellar@mail2.ietf.org>; Mon, 7 Jul 2025 09:28:36 -0700 (PDT)
Received: from mail-vk1-xa2a.google.com (mail-vk1-xa2a.google.com [IPv6:2607:f8b0:4864:20::a2a]) (using TLSv1.3 with cipher TLS_AES_128_GCM_SHA256 (128/128 bits) key-exchange X25519 server-signature ECDSA (P-256) server-digest SHA256) (No client certificate requested) by mail2.ietf.org (Postfix) with ESMTPS id D90E3401FD94 for <cellar@ietf.org>; Mon, 7 Jul 2025 09:27:44 -0700 (PDT)
Received: by mail-vk1-xa2a.google.com with SMTP id 71dfb90a1353d-5346b75d719so2799280e0c.1 for <cellar@ietf.org>; Mon, 07 Jul 2025 09:27:44 -0700 (PDT)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=or13.io; s=google; t=1751905664; x=1752510464; darn=ietf.org; h=cc:to:subject:message-id:date:from:in-reply-to:references :mime-version:from:to:cc:subject:date:message-id:reply-to; bh=UmBQhcWRwyeZ3ga6MuQFFN7DczMc+Xk4EothKmPTbCo=; b=RyqUGqEPYgZ5vYjnI88PZ/iJBN4rlBaPCG6YWo7qYh0/oGQorijRuqKZq1NBX6PXF6 Q49Q6JBLxr8uhSty/89B6OQbdLqZzDJicr9Wn/ZRILoe/CFy6txXHNLauCKQoOtzHcx3 np1Vv/NZwi7v10k34AG4QHQaWEgf/15lXbA6Q+i23VcK9GOj1WshNuFkOMgNpmNS4YQC y+vNOlWcsC2DWq/lXUsBhJKZT/MQeLQGKl9YRTIWADOAVfhCZ6fLvo12SmeKfROwzdQI L35WlhvUJaEzlOaLvC+GWHNFGl9445CDYohySkD4+CGGLwA65I2ID19bnVTRWEHXHeaf 6ITg==
X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1751905664; x=1752510464; h=cc:to:subject:message-id:date:from:in-reply-to:references :mime-version:x-gm-message-state:from:to:cc:subject:date:message-id :reply-to; bh=UmBQhcWRwyeZ3ga6MuQFFN7DczMc+Xk4EothKmPTbCo=; b=sjCANIU+ZeoZGVBk1YMWWDuntyi2yGJt8wQczIA/P7mHO24wD5OG2YbK0EKxVXKOi8 4rrAWA44nAa/p9vEStC0urvDT6vf3uQd/9of2cL8meTmaVOK36j5ElgmplDEFQIWQEAZ KDfXSvxKuiY8ca78Rr9g/APQeHERwFnHJzegw5aVVhVrxMUC+GWM505XIavxGtp/Vk9r TSUav5hgmGRYpTFONgPXOFVJ08IycyfrK/pcHjghmeefFUbdrXfhnzNt7ZOBy6vue7Ju iSOn5Zz1iKr0Ko7CCaPti5FAiabI6ulRcRdmGd9W+gpIMMEYXP4eQ00qD0SCLCQ00OhQ r+wA==
X-Gm-Message-State: AOJu0YziUP1cT1ZlFZvP2L+TzBI0Q6XH51IVywMMgxee8BzikMFLzGgM ZHgGqg4KZX3psAvCXtxXINjIkY8N5b0tWn7wD0VFIVboibWD+XXxk7NO91TdKR1G0wGMnKg2PG/ T9YTc3++qHh9lbZqWg1bTPSkHhvkFXAZVKijU/mIzfVkc72e2yfUnNtZFTA==
X-Gm-Gg: ASbGncvfPJZ71DFmB/xfYW078tl73ook2PMmJ9XFsOmNTpS8FLSWNT84Hlynsg/DQLi iO4OYG3jwaMmtuUloGMKMUkMMEU80DE1GODV+uvKvnralVm1CbvOZZE9Hrwy6BOfP4xdCCueuAk H1EwUhvLZ2OC8o9zccq93f99OcL8Y2kjqFTjjSCjsutmZilQvAyidw1wX6i5mkaRxJGRQk4bP7Z ioo5w==
X-Google-Smtp-Source: AGHT+IEjc/0nTQfckYh7CMYYhRcYPJMF65Yapy71jVMFmwATZOjICpcFqc0XyoXISZo8xDla9P33XpaAIcfDhLSkn8I=
X-Received: by 2002:a05:6122:3128:b0:535:aea0:795a with SMTP id 71dfb90a1353d-535c3788ec6mr335969e0c.1.1751905664092; Mon, 07 Jul 2025 09:27:44 -0700 (PDT)
MIME-Version: 1.0
References: <CAMzqgozwprvj=BFmhzmSR4oG=07xdOz=fh5hV6Z+ZmY_UKniSw@mail.gmail.com> <0468E722-3BEB-4C2C-82FC-0F2C6A6B0A1B@matroska.org>
In-Reply-To: <0468E722-3BEB-4C2C-82FC-0F2C6A6B0A1B@matroska.org>
From: Orie <orie@or13.io>
Date: Mon, 07 Jul 2025 11:27:33 -0500
X-Gm-Features: Ac12FXxpVax7tFId4_kPxpAh9SvSdHUNBAU0tnBuGcvWjkmx3FGHEfjm-ZrtJaw
Message-ID: <CAMzqgoxF7JdW6=18r1mzFbsPF6R0cYi-255FWHufy=hJp6bn7A@mail.gmail.com>
To: Steve Lhomme <slhomme@matroska.org>
Content-Type: multipart/alternative; boundary="000000000000c9e03c06395952e8"
Message-ID-Hash: 6UI2HTFXNU4HYUTPROLZGBW7WG2XZ3EX
X-Message-ID-Hash: 6UI2HTFXNU4HYUTPROLZGBW7WG2XZ3EX
X-MailFrom: orie@or13.io
X-Mailman-Rule-Misses: dmarc-mitigation; no-senders; approved; emergency; loop; banned-address; member-moderation; header-match-cellar.ietf.org-0; nonmember-moderation; administrivia; implicit-dest; max-recipients; max-size; news-moderation; no-subject; digests; suspicious-header
CC: cellar@ietf.org
X-Mailman-Version: 3.3.9rc6
Precedence: list
Subject: [Cellar] Re: AD Evaluation: draft-ietf-cellar-tags-18
List-Id: Codec Encoding for LossLess Archiving and Realtime transmission <cellar.ietf.org>
Archived-At: <https://mailarchive.ietf.org/arch/msg/cellar/xcnJXMtu-NzR3C7vkv9R7lvn2qQ>
List-Archive: <https://mailarchive.ietf.org/arch/browse/cellar>
List-Help: <mailto:cellar-request@ietf.org?subject=help>
List-Owner: <mailto:cellar-owner@ietf.org>
List-Post: <mailto:cellar@ietf.org>
List-Subscribe: <mailto:cellar-join@ietf.org>
List-Unsubscribe: <mailto:cellar-leave@ietf.org>
Hi Steve & Cellar, Inline replies prefixed with OS, sorry to have not gotten you this feedback in time for the draft cut off. Unless noted, you can assume I have no objection to your comments, but please point out if you are expecting a specific response from me. On Sun, Jun 29, 2025 at 5:55 AM Steve Lhomme <slhomme@matroska.org> wrote: > Hi Orie, > > Thanks a lot for your detailed review. I’ll add comments and some Merge > Requests with fixes where possible. > > On 20 Jun 2025, at 22:54, Orie <orie@or13.io> wrote: > > # Orie Steele, ART AD, comments for draft-ietf-cellar-tags-18 > CC @OR13 > > * line numbers: > - > https://author-tools.ietf.org/api/idnits?url=https://www.ietf.org/archive/id/draft-ietf-cellar-tags-18.txt&submitcheck=True > > * comment syntax: > - https://github.com/mnot/ietf-comments/blob/main/format.md > > ## Discuss > > ### Reduced version underspecified > > ``` > 225 TagString fields defined in this document with dates MUST have the > 226 following format: "YYYY-MM-DD hh:mm:ss.mss" or a reduced version. > ``` > > Suggest you provide ABNF for the reduced version, or provide format > strings for all valid reduced versions. > > > Done in https://github.com/ietf-wg-cellar/matroska-specification/pull/1019 > > I assume "YYYY-MM-DD hh" is a valid reduced version? > > > Yes it it. > > Having both ISO8601 and RFC3339 as informative references here is awkward. > Consider framing this "custom date format" as MUST be an RFC3339 date > string, without the T ... etc. > > > They are informative because our format differs from both and is never a > subset (AFAIK). > > Given the open nature of tags, I’m not sure we can impose a definite > format that MUST be used all the time. For example some new date/time tag > may decide to use exactly the same string format as found in another > container to avoid translations (possibly lossy). > OS: Your existing text is close to saying this, but you might avoid future comments by being even more direct, these datetime strings are not conformant to ISO8601 or RFC3339, but they appear similar. > > ### Digit grouping delimiters > > ``` > 240 TagString fields that require a floating-point number MUST use the > 241 "." mark instead of the "," mark. Only ASCII numbers "0" to "9" and > 242 the "." character MUST be used. The "." separator represents the > 243 boundary between the integer value and the decimal parts. If the > 244 string doesn't contain the "." separator, the value is an integer > 245 value. Digit grouping delimiters MUST NOT be used. > ``` > > What is a digit grouping delimiter? > > > I don’t think there’s a formal term for that in English, but it’s inspired > by https://en.wikipedia.org/wiki/Decimal_separator#Digit_grouping > > It’s the delimiter they talk about for the “digit grouping”. > > I assume "," is a digit grouping delimiter, are there others? > > > In French “,” is the decimal separator and “.” Is the digit grouping > delimiter. In English it’s the other way around. > > Are negative integers allowed? Do you mean to specify whole numbers and > positive decimal numbers instead? > > > It is not mentioned, but floating-point numbers can be negative. Should we > mention it ? > > There is no specification for decimal numbers, only the floating point > ones as strings are defined as they are subject (mis)interpretation. > OS: Not clear what the action is here, I suggest you provide ABNF for the strings that need to be recognized, to avoid confusion or subjective interpretation. > > ### UI MUST? > > ``` > 247 To display it differently for another locale, applications MUST > 248 support auto replacement on display. > ``` > > How is this implemented? > I don't understand this normative requirement. > Consider making this informative / motivational for the number formats > section. > > > Indeed, it cannot really be enforced nor is absolutely necessary. > I made it a recommendation in > https://github.com/ietf-wg-cellar/matroska-specification/pull/1020 > > ### Legacy character handling > > ``` > 250 In legacy media containers, it is possible that the "," character > 251 might have been used as a separator or that digit grouping delimiters > 252 might have been used. A Matroska Reader SHOULD consider the > 253 following character handling to parse such legacy formats: > > 255 * if multiple instances of the same non-number character are found, > 256 they are be ignored, > 257 * if only one "." character is found and no other non-number > 258 character is found, the "." is the integer-decimal separator, > 259 * if only one "," character is found and no other non-number > 260 character is found, the "," is a digit grouping delimiter, > 261 * any other non-number character is ignored. > ``` > > What is a legacy media container? (citation please) > > > These are Matroska files that been created a decade or two ago when the > tags specification wasn’t as defined as it is now. > > I’m not sure what a better expression would be (apart from what I wrote > above). We try to avoid the word “file” as it can also be a network stream. > > "SHOULD consider" ... what does this mean? > > > Readers do not absolutely have to “transform” the strings depending on > what they do with it. Since it’s formats not matching the specs, they could > be considered as invalid. > > These files may exist because of the old spec was under specified although > the normal interpretation would match what we have formally defined now. > > If these characters are possible to encounter, they ought to be allowed in > the specification, and recognized, if you are trying to specify a > transformation here, give a name to each side of the transformation and > describe how to handle any exceptions that might be encountered from > malformed data. > > > We want to have a simple format to parse, hence sticking with the more > common/assumed version. We don’t even know if files with non-compliant > strings. For example it’s highly unlikely there are multiple instances of > the same non-number character are found. > > This text should be considered as error handling of possibly bogus data. > OS: It could be clearer, if the policy is to raise an exception (fail) or simply ignore bogus data. If there is no guidance, expect software to implement whichever is least helpful : ) > > This section needs some work. > I suggest you focus on describing the numbers that exist, and then > recommend how to produce new numbers, only allowing for legacy formats as a > backward compatibility feature. > > > I feel that’s what we did, maybe not with the best wording. > > I would avoid describing transformations, unless you also describe > exception handling. > Is the inline replacement comment meant to imply that readers mutate the > data to display it? > > ## Comments > > ### synthetic data > > ``` > 125 * ARTIST = "Pet Shop Boys" > > 127 - LEAD_PERFORMER = "Neil Tennant" > > 129 o DATE_STARTED = "1981-08" > ``` > > Prefer to use synthetic data for examples. > In recent drafts, we have removed real artist names. > > > Is this a guideline we have to follow ? Because I feel using completely > made up artists will make it harder to understand what we are describing. > We even have non-normative links to make sure people can see what the real > world examples correspond to. > OS: This is just a comment, you may see similar feedback from the IESG. For this specification, real world examples seem helpful. > > Same comment for the later section: > > ``` > 416 * ARTIST = "Daft Punk" > > 418 * TITLE = "Da Funk" > > 420 * TOTAL_PARTS = "2" > ``` > > ### Guidance to DEs? > > ``` > 181 would be better for that. It's hard to define what should be in and > 182 what doesn't make sense in a file; thus, each demand needs to balance > 183 if it makes sense to be carried over in a file for storage and/or > 184 sharing or if it doesn't belong there. > ``` > > Is this section meant to be guidance to DEs? > Consider framing this as guidance to DEs. > > > What do you call DE ? > OS: Designated Experts, per: https://datatracker.ietf.org/doc/html/rfc8126#section-4.5 ... Please make sure to read this RFC thoroughly, IANA Considerations are an easy DISCUSS, if they are not clearly aligned with 8126. > Later: > > ``` > 186 We also need an official list simply for developers to be able to > 187 display relevant information in their own design, if they choose to > 188 support a list of meta-information they should know which tag has the > 189 wanted meaning so that other apps could understand the same meaning. > ``` > > This appears to be the value of maintaining a public registry, consider > framing this as motivation for the establishment and maintenance of the > registry, instead of motivation for tags > > > This is in a section called “Why Official Tags Matter”. This applies to > the whole document in general, so that includes the registry as well. It’s > referenced in multiple places in the document. We could add a reference in > the IANA section as well. > > Later: > > ``` > 1671 Matroska Tag Names for binary data are to be allocated according to > 1672 the "Specification Required" policy [RFC8126]. > ``` > > See https://datatracker.ietf.org/doc/html/rfc8126#section-4.6 > > """ > This policy is the same as Expert Review, with the additional > requirement of a formal public specification. > """ > > I would recommend organizing the guidance for designated experts into its > own section. > > > You mean grouping by First Come First Served and the ones that are > Specification Required, rather than tag types as they are now ? > OS: Similar to: https://datatracker.ietf.org/doc/html/rfc8126#section-4.12 > > ### UTF-8 "letters" > > ``` > 195 Official TagName values MUST consist of UTF-8 capital letters, > 196 numbers and the underscore character '_'. > > 198 Official TagName values MUST NOT contain any space. > > 200 Official TagName values MUST NOT start with the underscore character > 201 '_'; see Section 3.1. > ``` > > ABNF might be helpful for this. > > > That would definitely be useful. However there doesn’t seem to be any > logic for UTF-8 upper (or lower) case values, or even simply letters, as > opposed to symbols. > > Looking at RFC 5234 (ABNF) it doesn’t really take in account Unicode but > rather defined bits when it’s not ASCII (RFC 3629 - UTF-8 has a binary ABNF > definition). > > I can a name to describe a UTF-8 upper letter and then do the rest of the > ABFN with that. But I don’t think that’s valid. > OS: Consider if this is helpful: https://datatracker.ietf.org/doc/draft-bray-unichars/ (for defining the repertoire you are expecting software to recognize) > ### Why not MUST? > > ``` > 214 Multiple items SHOULD never be stored as a list in a single > 215 TagString. If there is more than one tag value with the same name to > 216 be stored, then more than one SimpleTag SHOULD be used. > ``` > > When can this suggestion be ignored? > > > That’s at least the case of files that were created before the tags > specification was refined. This is again a rule that was loosely defined in > the past. Most people probably interpreted it the way it is now defined but > it’s not 100% sure. > > On the other hand, as there’s no list delimiter defined anyway, lists > there were not properly split won’t be easily discovered. And for writers > we could enforce this rule from now on. > > Later: > > ``` > 218 Due to preexisting files where these formatting rules were not > 219 explicit, they are usually presented as rules that SHOULD be applied > 220 when possible, rather than MUST be applied at all times. It is > 221 RECOMMENDED to use strict formatting when writing new tag values. > ``` > > It's enough to say that "for backward compatibility, multiple items SHOULD > NOT...". > > And again, for new tag values, why not MUST? > > > I agree, we can do that. > > I turned the rule into a MUST in > https://github.com/ietf-wg-cellar/matroska-specification/pull/1021 > And added some explanation as to why it is easier said than done. > > ### Type UTF-8 > > ``` > 824 | SUBTITLE | UTF-8 | Sub Title of the entity. This is | > ``` > > Are emoji allowed? what about control characters? > > > Any UTF-8 character. This is stored in a binary format that doesn’t need > escaping. So any valid UTF-8 value is fine. > > Is there a more specific way to describe this type? > > Consider using https://datatracker.ietf.org/wg/precis/documents/ or > https://datatracker.ietf.org/doc/draft-bray-unichars/ . > > > Not sure what you are proposing here. The value is anything that is a > valid UTF-8 string. The meaning is a “sub title”. Maybe “under title” is > better or there are more appropriate word in English ? > It seems that “subtitle” is commonly used for that > https://www.writeitgreat.com/post/title-vs-subtitle-what-s-the-difference > > In the linked ID3 tag, the definition is “Subtitle/Description refinement”. > OS: See comment about repertoires above, please read the reference and consider if an attacker might abuse the ability to inject arbitrary unicode. > ### MAY exceed max depth > > ``` > 1642 memory of the host app by using very deep nesting. An host app MAY > 1643 add some limits to the amount of nesting possible to avoid such > 1644 issues. > ``` > > What does adding a limit mean here? > > Is everything below that depth ignored? > > Is it possible that adding such a limit could lead to other application > failure issues? > > > The limit is on the reader side and may be of multiple nature. I don’t > think we should go into details of all the potential implementation issues > and list all exhaustive ways of getting around them. It’s also up to the > reader if they want to consider the whole Tags element using too much > memory as invalid (for example via an exception in C++), or just stop > reading more tags when there’s no more memory. Some implementations may > consider 3 levels of nesting, some 123456 levels. > OS: Considering that without letting a limit each reader will reject different files, and there will not be good interop, I suggest discussing with the installation base and setting a reasonable limit, as a SHOULD... not a MAY. > > ### Why not must > > ``` > 1657 The Name corresponds to the value stored in the TagName element. > The > 1658 Name SHOULD always be written in all capital letters and contain no > 1659 space as defined in Section 3.2, > ``` > > Is this also to support backward compatibility? > > > Yes. And also because a lot of programs may just copy/paste the Name/Value > (key, value) from other formats right away in Matroska without trying to > map to a proper/official name. > > ## Nits > > ### in in > > ``` > 96 section, the different high level parts of Matroska as defined in in > 97 Section 4.5 of [RFC9559] and EBML Master Elements as defined in in > ``` > > > Fixed by > https://github.com/ietf-wg-cellar/matroska-specification/pull/1022 > > ### reads awk > > ``` > 178 (Section 6.1). This registry is not meant to have every possible > 179 information in a file. Matroska files are not meant to become a > ``` > > > Would replacing “information” (which is a vague quantity in English, > unlike French) with metadata work ? > > _______________________________________________ > Cellar mailing list -- cellar@ietf.org > To unsubscribe send an email to cellar-leave@ietf.org > > >
- [Cellar] AD Evaluation: draft-ietf-cellar-tags-18 Orie
- [Cellar] Re: AD Evaluation: draft-ietf-cellar-tag… Orie
- [Cellar] Re: AD Evaluation: draft-ietf-cellar-tag… Steve Lhomme
- [Cellar] Re: AD Evaluation: draft-ietf-cellar-tag… Orie
- [Cellar] Re: AD Evaluation: draft-ietf-cellar-tag… Spencer Dawkins at IETF
- [Cellar] Re: AD Evaluation: draft-ietf-cellar-tag… Steve Lhomme
- [Cellar] Re: AD Evaluation: draft-ietf-cellar-tag… Steve Lhomme
- [Cellar] Re: AD Evaluation: draft-ietf-cellar-tag… Spencer Dawkins at IETF
- [Cellar] Re: AD Evaluation: draft-ietf-cellar-tag… Steve Lhomme
- [Cellar] Re: AD Evaluation: draft-ietf-cellar-tag… Robert Sparks