Re: [apps-discuss] character repertoire for fragment identifiers

Sam Ruby <rubys@intertwingly.net> Sun, 11 January 2015 22:11 UTC

Return-Path: <rubys@intertwingly.net>
X-Original-To: apps-discuss@ietfa.amsl.com
Delivered-To: apps-discuss@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 5CD1F1A8820 for <apps-discuss@ietfa.amsl.com>; Sun, 11 Jan 2015 14:11:59 -0800 (PST)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.9
X-Spam-Level:
X-Spam-Status: No, score=-1.9 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, RCVD_IN_DNSWL_NONE=-0.0001] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id zb4CGarkHA5O for <apps-discuss@ietfa.amsl.com>; Sun, 11 Jan 2015 14:11:57 -0800 (PST)
Received: from cdptpa-oedge-vip.email.rr.com (cdptpa-outbound-snat.email.rr.com [107.14.166.227]) by ietfa.amsl.com (Postfix) with ESMTP id 4670E1A8824 for <apps-discuss@ietf.org>; Sun, 11 Jan 2015 14:11:53 -0800 (PST)
Received: from [98.27.51.253] ([98.27.51.253:5054] helo=rubix) by cdptpa-oedge03 (envelope-from <rubys@intertwingly.net>) (ecelerity 3.5.0.35861 r(Momo-dev:tip)) with ESMTP id 5B/3D-08411-825F2B45; Sun, 11 Jan 2015 22:11:52 +0000
Received: from [192.168.1.102] (unknown [192.168.1.102]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) (Authenticated sender: rubys) by rubix (Postfix) with ESMTPSA id 9101F140B53; Sun, 11 Jan 2015 17:11:52 -0500 (EST)
Message-ID: <54B2F527.7040404@intertwingly.net>
Date: Sun, 11 Jan 2015 17:11:51 -0500
From: Sam Ruby <rubys@intertwingly.net>
User-Agent: Mozilla/5.0 (X11; Linux x86_64; rv:31.0) Gecko/20100101 Thunderbird/31.3.0
MIME-Version: 1.0
To: Julian Reschke <julian.reschke@gmx.de>
References: <20140926010029.26660.82167.idtracker@ietfa.amsl.com> <DM2PR0201MB09602B351692D424A49C6B0DC3650@DM2PR0201MB0960.namprd02.prod.outlook.com> <CACweHNBN_Bv=jeXQ_VwXi2HzHKNEwZJ1NiF-BJJo_9-mhO60gQ@mail.gmail.com> <54A557E1.6050502@intertwingly.net> <CACweHNCQZg1U1u8U=-f6h0+BPnp6Wr_T=r_wGiPAbhTbuMCGWQ@mail.gmail.com> <54A94109.5010901@intertwingly.net> <00cf01d02cc7$d5dba4c0$4001a8c0@gateway.2wire.net> <54B16C2B.9050604@seantek.com> <54B17BBE.4000900@intertwingly.net> <54B18B61.8010308@seantek.com> <54B19435.8070401@intertwingly.net> <54B1B211.3050807@seantek.com> <54B1B682.3070609@intertwingly.net> <54B28E0F.8070306@gmx.de> <54B2936B.7030805@intertwingly.net> <05AD7DE2-1C54-45CD-B33A-13766D771E57@mnot.net> <54B2A2CD.5080502@gmx.de> <1A5BBD25-FEBD-49B1-9EFB-4EF8877BF0E7@mnot.net> <54B2A4F9.2070909@gmx.de> <54B2A894.4020201@intertwingly.net> <54B2ABA8.6030205@gmx.de> <54B2C6FA.80802@intertwingly.net> <54B2EE08.9040705@gmx.de>
In-Reply-To: <54B2EE08.9040705@gmx.de>
Content-Type: text/plain; charset="utf-8"; format="flowed"
Content-Transfer-Encoding: 7bit
X-RR-Connecting-IP: 107.14.168.142:25
X-Cloudmark-Score: 0
Archived-At: <http://mailarchive.ietf.org/arch/msg/apps-discuss/YE1Jd0M_P_l9J62iVnozXTdgeaQ>
Cc: apps-discuss@ietf.org
Subject: Re: [apps-discuss] character repertoire for fragment identifiers
X-BeenThere: apps-discuss@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: General discussion of application-layer protocols <apps-discuss.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/apps-discuss>, <mailto:apps-discuss-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/apps-discuss/>
List-Post: <mailto:apps-discuss@ietf.org>
List-Help: <mailto:apps-discuss-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/apps-discuss>, <mailto:apps-discuss-request@ietf.org?subject=subscribe>
X-List-Received-Date: Sun, 11 Jan 2015 22:11:59 -0000

On 01/11/2015 04:41 PM, Julian Reschke wrote:
> On 2015-01-11 19:54, Sam Ruby wrote:
>> ...
>> At some point, I would like to discuss what characters should be allowed
>> in a fragment, and would like RFC 3986 either to match the set we come
>> up with, or for RFC 3986 to be obsoleted by a specification that does.
>>
>> I've been proceeding by testing actual implementations and actually
>> talking to at least two authors of RFC 3986.  And, trust me, I plan to
>> have conversations with the third author.
>>
>> Each time I try to have this discussion, you keep returning to what RFC
>> 3986 currently says as of this moment.
>> ...
>
> I keep returning to URIs as things consisting of ASCII code points.

Indeed you do.  :-(

> If you want to talk about URI-like identifiers that can contain
> non-ASCII characters, we'll need to discuss RFC *3987*. As far as I can
> recall, there's more or less agreement that this needs to happen, but
> that's IMHO a different conversation than the one about RFC 3986.

Let me compare what you suggest should happen with what Roy Fielding 
suggests should happen.  For whatever reason, Roy's post didn't make the 
web archive, but my reply did, so I will Roy quote from:

https://lists.w3.org/Archives/Public/public-ietf-w3c/2014Dec/0088.html

 >>> The problem with RFC3987 was that it tried to define a new
 >>> addressing format instead of simply defining an arbitrary
 >>> reference and how to get from there to an interoperable URI. It
 >>> did not work because it wasn't written to handle arbitrary input
 >>> and could not keep up with changes in IDNA.

and

 > This does not mean there isn't value in coming up with a consistent
 > reference parsing model for HTML that reflects the browser
 > environment and results in a consistent URL DOM and address display.
 > It simply doesn't change what is in RFC3986, nor does it escape the
 > fact that a browser is still dependent on RFC3986 to produce a valid
 > URI out of whatever it happens to parse when it eventually chooses to
 > use that URI in an IETF protocol like HTTP.

I'd like to explore what Roy is suggesting: layering URLs on URIs, 
without going through a URL layered on IRI layered on URI approach. 
Doing so may require changes to both the definition of URLs and URIs.

In particular, I'd like to see how close we can come to a goal where the 
output of URL parsing followed by URL stringification is always a valid URI.

There is no question that a change to the "ASCII-ness" of schemes is 
clearly a non-starter.  But I would hope that we could discuss whether a 
change to the "ASCII-ness" of fragments is a possibility.  In 
particular, I don't believe that there is any reason why a URL parser 
(in a browser or elsewhere) needs to percent encode fragments.

It may come to pass that Roy is the only person who shares the goals 
that he described.  In which case, the end result will be an appendix to 
the URL standard which describes the real world use cases that RFC 3986 
does not address which prevent a proper layering.  Additionally, there 
would be a description of what the set of outputs from a URL 
parse/stringification process does produce, i.e., a description of what 
RFC 3986 should have said.

- Sam Ruby