Re: [tcpm] AD Review of draft-ietf-tcpm-hystartplusplus-09

Martin Duke <martin.h.duke@gmail.com> Thu, 08 September 2022 17:05 UTC

Return-Path: <martin.h.duke@gmail.com>
X-Original-To: tcpm@ietfa.amsl.com
Delivered-To: tcpm@ietfa.amsl.com
Received: from localhost (localhost [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 156CCC1524D6; Thu, 8 Sep 2022 10:05:16 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -2.106
X-Spam-Level:
X-Spam-Status: No, score=-2.106 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, FREEMAIL_FROM=0.001, HTML_MESSAGE=0.001, RCVD_IN_DNSWL_BLOCKED=0.001, RCVD_IN_ZEN_BLOCKED_OPENDNS=0.001, SPF_HELO_NONE=0.001, SPF_PASS=-0.001, T_SCC_BODY_TEXT_LINE=-0.01] autolearn=ham autolearn_force=no
Authentication-Results: ietfa.amsl.com (amavisd-new); dkim=pass (2048-bit key) header.d=gmail.com
Received: from mail.ietf.org ([50.223.129.194]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id fqiyfGXQYmU1; Thu, 8 Sep 2022 10:05:15 -0700 (PDT)
Received: from mail-qv1-xf36.google.com (mail-qv1-xf36.google.com [IPv6:2607:f8b0:4864:20::f36]) (using TLSv1.3 with cipher TLS_AES_128_GCM_SHA256 (128/128 bits) key-exchange X25519 server-signature RSA-PSS (2048 bits) server-digest SHA256) (No client certificate requested) by ietfa.amsl.com (Postfix) with ESMTPS id 19AA9C1524B1; Thu, 8 Sep 2022 10:05:15 -0700 (PDT)
Received: by mail-qv1-xf36.google.com with SMTP id v15so10973016qvi.11; Thu, 08 Sep 2022 10:05:15 -0700 (PDT)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20210112; h=cc:to:subject:message-id:date:from:in-reply-to:references :mime-version:from:to:cc:subject:date; bh=HeKoflHjMYrLU1DrReNr66py7VWdNHm4KdrH0HjY8GE=; b=i3OtazgaZf316FzOGkQjrc6J3I4eAJSCbIbtMgJqraMAw+jITfcFkIP1EyG4kfCzbk V9Eq9u2Fiv37PhesL8FOL3WrI6WwfeRYGdRsu+TVfSJ1S7CJlM4w9+id09b4m487vWlx wDqDw9ARqBbTpZ3ZUWLt0TpsZlG1aX38ZxCF2B9JG8dsANJC03h9O/EXe3cTC2V6DMFZ OG5miuxzYSRBAxl5vlKFp7CkkcVyhbTW8yghsFHLKEDlqYFoYMLHuY6Uh5ELa3mQL2+p 04Xd7QgOlvQC9wkEeJn7NJFeO40PwIhXSSeRzLLeIIzi5AkzNJ8VAy5N5+76FdhySdEy e3zA==
X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20210112; h=cc:to:subject:message-id:date:from:in-reply-to:references :mime-version:x-gm-message-state:from:to:cc:subject:date; bh=HeKoflHjMYrLU1DrReNr66py7VWdNHm4KdrH0HjY8GE=; b=3CJ4MBszYbtHZo3XueszW4SQ4PUqsJ8P456fJZGd0+Wq+rI3z0Wm81dgr16shW+MGO EdEwSUx9yGM7HoSHR2Smw81cbFaUom7ylafa471885OM3FDI1A7/rt6oH4p6WaTBlT+u EXfMDDeoVLHMaQaQ08dcy95DP/V1QvdeK9Ic1zJ5p3Bh5MJAtxMosimvu05JYZts8w0u hJ2nku3NLOtkkVWrwuZp+xihDK9o+ASwAM8lyhO+HMo2L4x2T/hmBGmrafrZKnYtmzJV vUyfpvS3OSnap9pW8CgoEkz+kLrhlHAfSrRlkqQE5LQQCXstypEn1UgzbG32SYNMhsob t3OQ==
X-Gm-Message-State: ACgBeo2mFAsvMMg+ILJQbpIl+kQuHco3NzEgfEDYYhCPbeYLMB69oA3A +liPqpLsJ23IV1MA4ovUlmGPxZVgvyGLMWceDQQ=
X-Google-Smtp-Source: AA6agR7tuQNtIaw2V9zIgaBgPaw3q6/T1lYtymXpjoXtPor9A3NDzSpxkUavYhvndF9s7cPqwewdPkICPOj29YmcRYM=
X-Received: by 2002:a0c:8d0a:0:b0:4ac:82a6:7a9a with SMTP id r10-20020a0c8d0a000000b004ac82a67a9amr1693099qvb.61.1662656713611; Thu, 08 Sep 2022 10:05:13 -0700 (PDT)
MIME-Version: 1.0
References: <CAM4esxTikRRRLOtmO4bezXvjjDiQ3cpRNqtT_2YaEUQrFUrECw@mail.gmail.com> <0E94A985-516C-4287-9789-50D3A682211B@fh-muenster.de> <CAM4esxS-3kZDLp-1n3JM3jOH=4ZLNaf0vsszMgqrJ7ss=jxcCg@mail.gmail.com> <D98D5BA8-B5B8-42F7-AF61-352235D8BC39@fh-muenster.de> <CAM4esxQNfuMKq+pj0v2+mO8Mxq8nbpgmVDai-WgLvgCfXrtibQ@mail.gmail.com> <741E2E66-0EC9-48E3-BF8E-4EFE000567D2@fh-muenster.de> <CADVnQym61GQ5zqE8+VpkB9=D03UYwWyVodbJECY319vkP9qDjQ@mail.gmail.com> <CAM4esxSobvx4yiYkGe0xgPEO7=s_tCo4-ieWHNedai+xmeZiOA@mail.gmail.com> <CADVnQykeDYjp=wGpGByU2dQG-9r8s6OxazodJAp9tcBR3-2YDw@mail.gmail.com>
In-Reply-To: <CADVnQykeDYjp=wGpGByU2dQG-9r8s6OxazodJAp9tcBR3-2YDw@mail.gmail.com>
From: Martin Duke <martin.h.duke@gmail.com>
Date: Thu, 08 Sep 2022 10:05:02 -0700
Message-ID: <CAM4esxTFtjUDAumeT+CST3wq8C4A-J3hN-qySWRN7LLzV04Gog@mail.gmail.com>
To: Neal Cardwell <ncardwell@google.com>
Cc: Michael Tuexen <tuexen@fh-muenster.de>, draft-ietf-tcpm-hystartplusplus.all@ietf.org, "tcpm@ietf.org Extensions" <tcpm@ietf.org>, Yuchung Cheng <ycheng@google.com>
Content-Type: multipart/alternative; boundary="000000000000cc5da205e82d6e71"
Archived-At: <https://mailarchive.ietf.org/arch/msg/tcpm/M9vQ2HMt0jDEbf1hYyYvi2C6KCU>
Subject: Re: [tcpm] AD Review of draft-ietf-tcpm-hystartplusplus-09
X-BeenThere: tcpm@ietf.org
X-Mailman-Version: 2.1.39
Precedence: list
List-Id: TCP Maintenance and Minor Extensions Working Group <tcpm.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/tcpm>, <mailto:tcpm-request@ietf.org?subject=unsubscribe>
List-Archive: <https://mailarchive.ietf.org/arch/browse/tcpm/>
List-Post: <mailto:tcpm@ietf.org>
List-Help: <mailto:tcpm-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/tcpm>, <mailto:tcpm-request@ietf.org?subject=subscribe>
X-List-Received-Date: Thu, 08 Sep 2022 17:05:16 -0000

To be clear, the non-normative experiment report can also say "we were also
running non-standard ABC with L=8", but as it stands the normative part
can't have a reference to L that requires 3465 to fully understand

On Thu, Sep 8, 2022, 10:00 Neal Cardwell <ncardwell@google.com> wrote:

> On Thu, Sep 8, 2022 at 12:54 PM Martin Duke <martin.h.duke@gmail.com>
> wrote:
>
>> Hi Neal, thanks for the data.
>>
>> Given its wide deployment, I'm happy for the spec to allow L_max =
>> infinity, assuming the community is fine with it.
>>
>> I don't care that much, in theory, whether L limits are defined in this
>> document or in another one. However, in practice adding about 2 paragraphs
>> to this document would allow us to obsolete 3465. If we can rapidly reach
>> consensus on this, it seems like the expeditious thing to do.
>>
>
> Thanks, Martin. Adding about 2 paragraphs to this document in order to
> allow us to obsolete 3465 sounds great to me. :-)
>
> The alternative is to simply eliminate L from Hystart++, and we "really
>> know" what people will do in production.
>>
>
> Eliminating L from Hystart++ sounds like a nice simplification. But there
> is also virtue in Microsoft documenting what is implemented and tested. So
> I don't have strong feelings about whether L is kept in Hystart++ or not,
> and I hope we hold up the Hystart++ draft due to issues surrounding L. :-)
>
> regards,
> neal
>
>
>
>>
>> On Thu, Sep 8, 2022 at 9:42 AM Neal Cardwell <ncardwell@google.com>
>> wrote:
>>
>>>
>>>
>>> On Wed, Sep 7, 2022 at 5:46 PM <tuexen@fh-muenster.de> wrote:
>>>
>>>> > On 7. Sep 2022, at 22:44, Martin Duke <martin.h.duke@gmail.com>
>>>> wrote:
>>>> >
>>>> >
>>>> >
>>>> > On Wed, Sep 7, 2022 at 1:15 PM <tuexen@fh-muenster.de> wrote:
>>>> > > On 7. Sep 2022, at 21:40, Martin Duke <martin.h.duke@gmail.com>
>>>> wrote:
>>>> > >
>>>> > > I would be happier with some sort of limit on L. Maybe: the
>>>> positive integer L SHOULD be 2 and MUST be no more than 8(?) But whatever
>>>> numbers the community is comfortable with are fine with me.
>>>> > Hi Martin,
>>>> >
>>>> > the crucial point here is that we know that MS used L = 8. But we
>>>> don't have any information,
>>>> > whether other values of L are worse or better, what the tradeoffs are
>>>> and in particular
>>>> > what traffic patterns impact the choice of L. Wouldn't we need such
>>>> an analysis / results
>>>> > from an experiment for selecting an upper limit L_MAX? I would assume
>>>> that it might
>>>> > have an impact whether pacing is used or not.
>>>> >
>>>> > Well the community has at least some data that L = 8 is safe on the
>>>> internet. Maybe some other practitioners have data for 16 or 32 or
>>>> whatever. It would be reasonable to set some sort of limit based on the
>>>> limits of our empirical limit.
>>>> Sure. My point was just that the results provided my MS show that L_MAX
>>>> <= 8 seems to
>>>> be safe in the scenario they looked at. But there is no statement that
>>>> L_MAX > 8 is
>>>> bad...
>>>> >
>>>> > It would be fine to say "one SHOULD NOT exceed L = foo if you're not
>>>> pacing"
>>>> >
>>>> >
>>>> > >
>>>> > > I'd also be happier with a cut-and-paste of those considerations
>>>> from 3465, which should not be a lot of work.
>>>> > >
>>>> > > The lower-effort approach is just to strike L from the document and
>>>> let people do what they've doing, which is use Ls well in excess of 2. But
>>>> I'd prefer we'd actually write it down because
>>>> > > (1) our standards are better if they reflect widespread behavior in
>>>> the internet; and
>>>> > I agree with the above.
>>>> > > (2) it means we can finally make 3465 historic, as we have
>>>> standards-track documents that cover virtually all of its material.
>>>> > >
>>>> > > If the community really cannot converge on sensible limits for L,
>>>> then I'd rather publish with no L than wait for a deadlock to resolve.
>>>> > In my view, the selection of L is not part of the core of the
>>>> document.
>>>> >
>>>> > This document standardizes a new slow start algorithm, which includes
>>>> L as a parameter, for which there is no standard.
>>>> My point was that there is no L_MAX given here.
>>>> >
>>>> > >
>>>> > > But as a way forward, I'd suggest:
>>>> > > 1) the authors propose text that sets bounds on L (a SHOULD and a
>>>> MUST) and cut-and-pastes the considerations.
>>>> > Aren't the consideration in RFC 3465 coming to the conclusion: L
>>>> SHOULD be 1, MAY be 2, MUST NOT be larger than 2?
>>>> > How to use the same considerations to come to a different conclusion:
>>>> L SHOULD be 2, MUST NOT be larger than L_MAX
>>>> > for L_MAX > 2?
>>>> >
>>>> > Re: SHOULD 1 or 2, I don't have a strong opinion but I feel like the
>>>> 3465 experiment shows that 2 is safe, and somewhat beneficial given the
>>>> prevalence of ack thinning? If there is no consensus for this, I'm happy to
>>>> go with 1.
>>>> >
>>>> > L_MAX: My preference is to set a value that is widely deployed with
>>>> good data that it is safe to do so
>>>> Let us see if people provide data which values are good and which
>>>> values are bad...
>>>>
>>>
>>> FWIW, Linux TCP has been using L_MAX=infinity since 2013:
>>>
>>>
>>> https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=9f9843a751d0a2
>>>
>>> IMHO trying to limit bursts with an L parameter is misguided:
>>>
>>> + AN L limit is only a partial solution; TCP RFCs still allow massive
>>> line-rate bursts of cwnd when restarting from idle, which is super-common
>>> in some very common Internet workloads: web traffic, streaming video, RPC.
>>>
>>> + To avoid such bursts, whether restarting from idle or not, TCP stacks
>>> should enable pacing.
>>>
>>> + Once pacing is enabled, having an L limit purely shoots your flow in
>>> the foot, causing it to fail to double cwnd each round trip in slow start
>>> in the very common case of the many ubiquitous aggregation mechanisms that
>>> cause ACKs to arrive for dozens of packets at a time (e.g.. Linux GRO/LRO
>>> commonly aggregate 45 packets into a single unit that is processed and
>>> ACKed as a unit).
>>>
>>> IMHO ideally burst control, pacing, and L limits should be separated
>>> from Hystart++, and should have their own RFC.
>>>
>>> And whatever is done, at a minimum Hystart++ should not depend on
>>> RFC3465 (ABC) (though it could mention it in passing to provide context).
>>>
>>> my two cents,
>>> neal
>>>
>>>