[art] Re: draft-bray-unichars-09 review

Tim Bray <tbray@textuality.com> Tue, 08 October 2024 16:51 UTC

Return-Path: <tbray@textuality.com>
X-Original-To: art@ietfa.amsl.com
Delivered-To: art@ietfa.amsl.com
Received: from localhost (localhost [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 8A94CC1CAF42 for <art@ietfa.amsl.com>; Tue, 8 Oct 2024 09:51:14 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -2.107
X-Spam-Level:
X-Spam-Status: No, score=-2.107 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, HTML_MESSAGE=0.001, SPF_HELO_NONE=0.001, SPF_PASS=-0.001, T_SCC_BODY_TEXT_LINE=-0.01, URIBL_DBL_BLOCKED_OPENDNS=0.001, URIBL_ZEN_BLOCKED_OPENDNS=0.001] autolearn=ham autolearn_force=no
Authentication-Results: ietfa.amsl.com (amavisd-new); dkim=pass (1024-bit key) header.d=textuality.com
Received: from mail.ietf.org ([50.223.129.194]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id Jrxj3XPvlc3M for <art@ietfa.amsl.com>; Tue, 8 Oct 2024 09:51:10 -0700 (PDT)
Received: from mail-pj1-x102c.google.com (mail-pj1-x102c.google.com [IPv6:2607:f8b0:4864:20::102c]) (using TLSv1.3 with cipher TLS_AES_128_GCM_SHA256 (128/128 bits) key-exchange X25519 server-signature ECDSA (P-256) server-digest SHA256) (No client certificate requested) by ietfa.amsl.com (Postfix) with ESMTPS id 6B3BBC1840CC for <art@ietf.org>; Tue, 8 Oct 2024 09:51:10 -0700 (PDT)
Received: by mail-pj1-x102c.google.com with SMTP id 98e67ed59e1d1-2e188185365so4808434a91.1 for <art@ietf.org>; Tue, 08 Oct 2024 09:51:10 -0700 (PDT)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=textuality.com; s=google; t=1728406270; x=1729011070; darn=ietf.org; h=cc:to:subject:message-id:date:from:in-reply-to:references :mime-version:from:to:cc:subject:date:message-id:reply-to; bh=g3SpTXZpz2QzGuyWA562On+F6MatFDBlt+/O2mZqWZU=; b=boCC2D1qZXuhW+BHX+aSjClN1fcENJmp5gEQxfJX6xMRUppybK0RpKF85IvjCMoDv0 bV92TTtRu0oou68z+GGb2Oz1Youbye/3EXwoujVsjFuGRH6ALcigWWl3siBeWVo7wLPk RkSyzhjD5hD3SMANVNz/q8pU9pKnkGdbXTtLU=
X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1728406270; x=1729011070; h=cc:to:subject:message-id:date:from:in-reply-to:references :mime-version:x-gm-message-state:from:to:cc:subject:date:message-id :reply-to; bh=g3SpTXZpz2QzGuyWA562On+F6MatFDBlt+/O2mZqWZU=; b=EadbBuMBcOZdkaPf8MLxqwMuR26pkftANIWvmUaXvhUEX4Ld4+LJW35rNj4pj+2dr8 G5Coxg7xxBBLuZ2Y9aHx6AoEtDa39qct4+ESAejRbqJ/TADs3LBR/5TUe4CEL03z/BU4 KYmv0tJ8zoOis3yorPRiK6Ro8Q/QTE+l5DlQX2Ls9vQo0X9WHQyBZdOAFSloqDo2p+Xo nfPUtKL9qxSlw2g0ROQcFuXL8i87LIuWnuGF7h/r8fm9TF2qijhM7I2szlvRw03tfYra J8EFE9dkX+3PzR+ESgRQse61mZZd/bQKtoYDCHuwLtKe88Lhfb0wCnncrX44aF8oMuab Tc4A==
X-Forwarded-Encrypted: i=1; AJvYcCVlmyFOIQXURlCN86x/dfzImjOmC+bPgZXX1ov4FJo6IvWbmPD9trwEpVcnPsRu7WOCVRo=@ietf.org
X-Gm-Message-State: AOJu0YwL/28qzkymeksGr8Y1UCNnY7EXP/R6APgnRqpI0kuJCjSE41i6 ylLdqE8CcnMvhuY1GiWCYF1pbNBjUPexFw5Nfm27hshBNpsZuQjefKw50JSHL/zNt6ZapzRwTf5 +gA9ghXucFbAwJt/HodjPc1Dz812O785LjaJBnHPdVwCaQYf9
X-Google-Smtp-Source: AGHT+IFSIkm8LvJoWWzS2ImvNN1YzkE9LuWPh1BMwe0FK6Gjuq2JnWrCE4Z5O9v1RTScG3azTkfUeR20SfdPhEDcFWY=
X-Received: by 2002:a17:90a:9512:b0:2d8:9c97:3c33 with SMTP id 98e67ed59e1d1-2e1e6354187mr19682088a91.28.1728406269738; Tue, 08 Oct 2024 09:51:09 -0700 (PDT)
Received: from 1064022179695 named unknown by gmailapi.google.com with HTTPREST; Tue, 8 Oct 2024 09:51:08 -0700
Received: from 1064022179695 named unknown by gmailapi.google.com with HTTPREST; Tue, 8 Oct 2024 09:51:04 -0700
MIME-Version: 1.0 (Mimestream 1.4.1)
References: <CAN8C-_JkHcer1MbvwSrtFxFu_38JRy-_JstYOmAMdOVCdCe-0g@mail.gmail.com>
In-Reply-To: <CAN8C-_JkHcer1MbvwSrtFxFu_38JRy-_JstYOmAMdOVCdCe-0g@mail.gmail.com>
From: Tim Bray <tbray@textuality.com>
Date: Tue, 08 Oct 2024 09:51:08 -0700
Message-ID: <CAHBU6itokOYJb6ZWUfvct7O7AqDsULvYUrtBqm25Wy-7TxOM4Q@mail.gmail.com>
To: Orie Steele <orie@transmute.industries>
Content-Type: multipart/alternative; boundary="000000000000bc74350623f9f13c"
Message-ID-Hash: GJ4342LGJRQVDBEQHH7GH3CTMQMZOLA2
X-Message-ID-Hash: GJ4342LGJRQVDBEQHH7GH3CTMQMZOLA2
X-MailFrom: tbray@textuality.com
X-Mailman-Rule-Misses: dmarc-mitigation; no-senders; approved; emergency; loop; banned-address; member-moderation; header-match-art.ietf.org-0; nonmember-moderation; administrivia; implicit-dest; max-recipients; max-size; news-moderation; no-subject; digests; suspicious-header
CC: draft-bray-unichars@ietf.org, draft-bray-unichars.shepherd@ietf.org, art@ietf.org
X-Mailman-Version: 3.3.9rc5
Precedence: list
Subject: [art] Re: draft-bray-unichars-09 review
List-Id: Applications and Real-Time Area Discussion <art.ietf.org>
Archived-At: <https://mailarchive.ietf.org/arch/msg/art/zF4Kc7cVn9OwYEAzBpAklp5VXys>
List-Archive: <https://mailarchive.ietf.org/arch/browse/art>
List-Help: <mailto:art-request@ietf.org?subject=help>
List-Owner: <mailto:art-owner@ietf.org>
List-Post: <mailto:art@ietf.org>
List-Subscribe: <mailto:art-join@ietf.org>
List-Unsubscribe: <mailto:art-leave@ietf.org>

On Oct 4, 2024 at 7:22:12 AM, Orie Steele <orie@transmute.industries> wrote:

>
> ## Comments
>
> ### Lack of BCP14 Guidance
>

[This is just me, haven’t talked over Orie’s contributions with Paul yet.]

This is a really interesting issue.  Our vision of how this would be used
successfully was along the lines of:


   1. WG takes up a draft. That draft says "the value of this field is
   Unicode text." Maybe it even says "…encoded in UTF-8.”
   2. Someone (co-chair, WG member, AD, last-call commenter) says “Which
   Unicode characters? Go look at [Unichars].”
   3. The draft editor says “Oh, OK” and specifies one of the subsets from
   [Unichars] or maybe [PRECIS].


So, how can we use BCP14 to encourage this happening?  What would people’s
feelings be if we added a sentence, perhaps as a new paragraph in the
section 2 introduction, just before 2.1, reading as follows:

“It is RECOMMENDED that designers of protocols and data formats, for any
data field which contains textual data, consider the issues discussed in
this document. They SHOULD specify, for that data field, the use of one of
the subsets specified in this document or one of the profiles specified in
[PRECIS].”

I’ve never written BCP14 language that operates at the meta level like
this, aimed at designers of new protocols and data formats, rather than
describing a specific protocol/format. Is it even a thing? Are there other
examples?

### Recommendation to allow unassigned code points?
>
> ```
> 125   than 150,000 have been assigned to characters.  It is difficult to
> 126   specify that unassigned code points should be avoided, because they
> 127   regularly become assigned when new characters are added to Unicode.
> ```
>
> This is bordering on a recommendation, should it be strengthened to BCP14
> language?
>

Once again, I’m a little nervous about BCP14 at the meta level. That aside,
what about language like

“Designers of protocols and data formats SHOULD NOT attempt to forbid the
use of code points on the basis that they are unassigned at the time the
specification is created."

### What about CBOR or protocol buffers?
>
> ```
> 145   a single transformation format.  UTF-8 is widely used for
> 146   interoperable data formats such as JSON, YAML, and XML.
> ```
>

Hmm. Nothing against CBOR but the list could get pretty long with Protobufs
and Thrift and Avro etc etc. I want to avoid giving the impression that
this list is exhaustive rather than suggestive. Anyone got editorial
suggestions.

### Where is UTF-16 used?
>
> ```
> 178   A surrogate which occurs in text encoded in any transformation format
> 179   other than UTF-16 has no meaning and may cause malfunction in
> 180   software that encounters it.  In particular, it is impossible to
> 181   represent a surrogate in well-formed UTF-8.
> ```
>
> Some guidance here on where to expect UTF-16 could help explain where to
> expect surrogates.
>

UTF-16 is essentially never used on the wire. I for one have never seen it
in my decades of work in this space.  Does anyone have a counter-example?
Or even better, does anyone know of an RFC that says “you MUST NOT put
UTF-16 on the wire.”? Because thankfully, nobody does.

### Where are legacy codes still used in IETF protocols?
>
> ```
> 198   Aside from the useful controls, the control codes are mostly obsolete
> 199   and generally lack interoperable semantics.  This document uses the
> ```
>
> It would be good to call out any known places where these codes still have
> interoperable semantics.
>

I agree. I wouldn’t know where to look to find them. Does anyone have any
suggestions?

### Citation for this?
>
> ```
> 232   Surrogate code points have been observed to cause software failures.
> ```
>
Ideally in an RFC.
>

Sadly, I guess “Tim’s bitter personal experiences with legacy Java and C
code behind BigTech firewalls” doesn’t qualify.  I agree with Orie’s
implication that if we can’t back this up with a good reference we probably
shouldn’t say it.

### Clearer guidance on handling problematic input
>
> ```
> 260   Reasonable options for dealing with problematic input include, first,
> 261   rejecting text containing problematic code points, and second,
> 262   replacing them with placeholders.  (As an exception, [UNICODE] notes
> 263   that it may in some cases be appropriate, specifically for
> 264   noncharacters, to treat them as non-problematic unassigned code
> 265   points.)
> ```
>
> This could be framed with BCP14 language... The phrasing along with the
> parenthetical make this slightly hard to understand.
>

I was going to say we should strengthen the callout to [RFC9413] but I just
checked and, while it is very good, it also has no BCP14 language. I’d like
to say something like “Designers MUST consult [RFC9413] and [UNICODE]” but
that’s sort of meta-meta. I agree, the language about [UNICODE] should be
pulled out of parantheses and the reference should specify the [UNICODE]
section.

Oh, and thanks Orie, those are useful remarks.

 -T