Re: [tcpm] WGLC for draft-ietf-tcpm-alternativebackoff-ecn-06

Michael Welzl <michawe@ifi.uio.no> Sat, 17 March 2018 07:17 UTC

Return-Path: <michawe@ifi.uio.no>
X-Original-To: tcpm@ietfa.amsl.com
Delivered-To: tcpm@ietfa.amsl.com
Received: from localhost (localhost [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 8DF6A127076 for <tcpm@ietfa.amsl.com>; Sat, 17 Mar 2018 00:17:24 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -4.21
X-Spam-Level:
X-Spam-Status: No, score=-4.21 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, RCVD_IN_DNSWL_MED=-2.3, SPF_PASS=-0.001, T_RP_MATCHES_RCVD=-0.01, URIBL_BLOCKED=0.001] autolearn=ham autolearn_force=no
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id K8VjeenlZgr1 for <tcpm@ietfa.amsl.com>; Sat, 17 Mar 2018 00:17:22 -0700 (PDT)
Received: from mail-out02.uio.no (mail-out02.uio.no [IPv6:2001:700:100:8210::71]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by ietfa.amsl.com (Postfix) with ESMTPS id 44D23124BFA for <tcpm@ietf.org>; Sat, 17 Mar 2018 00:17:22 -0700 (PDT)
Received: from mail-mx04.uio.no ([129.240.10.25]) by mail-out02.uio.no with esmtps (TLSv1.2:ECDHE-RSA-AES256-GCM-SHA384:256) (Exim 4.90_1) (envelope-from <michawe@ifi.uio.no>) id 1ex65d-000A93-Mm; Sat, 17 Mar 2018 08:17:17 +0100
Received: from 93-58-133-64.ip158.fastwebnet.it ([93.58.133.64] helo=[10.0.0.2]) by mail-mx04.uio.no with esmtpsa (TLSv1.2:ECDHE-RSA-AES256-GCM-SHA384:256) user michawe (Exim 4.90_1) (envelope-from <michawe@ifi.uio.no>) id 1ex65b-0003af-SD; Sat, 17 Mar 2018 08:17:17 +0100
Content-Type: text/plain; charset="utf-8"
Mime-Version: 1.0 (Mac OS X Mail 10.3 \(3273\))
From: Michael Welzl <michawe@ifi.uio.no>
In-Reply-To: <alpine.DEB.2.20.1803170143270.3710@dx6-cs-02.pc.helsinki.fi>
Date: Sat, 17 Mar 2018 08:17:12 +0100
Cc: Michael Tuexen <tuexen@fh-muenster.de>, tcpm@ietf.org
Content-Transfer-Encoding: quoted-printable
Message-Id: <F5C780AF-4D6C-46F7-816C-0B540606A465@ifi.uio.no>
References: <7A3A6ECE-550B-4E5F-9D61-83C8969A7B93@fh-muenster.de> <alpine.DEB.2.20.1803170143270.3710@dx6-cs-02.pc.helsinki.fi>
To: Markku Kojo <kojo@cs.helsinki.fi>
X-Mailer: Apple Mail (2.3273)
X-UiO-SPF-Received: Received-SPF: neutral (mail-mx04.uio.no: 93.58.133.64 is neither permitted nor denied by domain of ifi.uio.no) client-ip=93.58.133.64; envelope-from=michawe@ifi.uio.no; helo=[10.0.0.2];
X-UiO-Spam-info: not spam, SpamAssassin (score=-4.9, required=5.0, autolearn=disabled, AWL=0.075, TVD_RCVD_IP=0.001, UIO_MAIL_IS_INTERNAL=-5, uiobl=NO, uiouri=NO)
X-UiO-Scanned: E94061EA2356C5D7569B60083E0E2EF7211C8579
Archived-At: <https://mailarchive.ietf.org/arch/msg/tcpm/s7S0NAaSmRB5vN9SCUJXYBQ9PJY>
Subject: Re: [tcpm] WGLC for draft-ietf-tcpm-alternativebackoff-ecn-06
X-BeenThere: tcpm@ietf.org
X-Mailman-Version: 2.1.22
Precedence: list
List-Id: TCP Maintenance and Minor Extensions Working Group <tcpm.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/tcpm>, <mailto:tcpm-request@ietf.org?subject=unsubscribe>
List-Archive: <https://mailarchive.ietf.org/arch/browse/tcpm/>
List-Post: <mailto:tcpm@ietf.org>
List-Help: <mailto:tcpm-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/tcpm>, <mailto:tcpm-request@ietf.org?subject=subscribe>
X-List-Received-Date: Sat, 17 Mar 2018 07:17:25 -0000

Dear Markku,

Thank you very much for reading this!
Answers in line below:


> On Mar 17, 2018, at 1:52 AM, Markku Kojo <kojo@cs.helsinki.fi> wrote:
> 
> Hi all,
> 
> Maybe slightly late wg last call review of draft-ietf-tcpm-alternativebackoff-ecn-06.
> 
> The draft is mostly well written and clear, and it is also improved since version -00 or -01 I last read it.
> 
> I believe the intention of the draft is just to allow and propose experimentation using a larger multiplicative decrease factor when responding to ECN indication than when responding to packet loss.
> 
> There is, however, a significant change that seem to have occured from version -04 to -05 that does much more than this and is therefore crucial to address properly. The text and reads now in -06 as follows:
> 
> 
>  It updates the following text in section 6.1.2 of the ECN
>  specification [RFC3168] :
> 
>     The indication of congestion should be treated just as a
>     congestion loss in non-ECN-Capable TCP.  That is, the TCP source
>     halves the congestion window "cwnd" and reduces the slow start
>     threshold "ssthresh".
> 
>  Replacing this with:
> 
>     Receipt of a packet with the ECN-Echo flag SHOULD trigger the TCP
>     source to set the slow start threshold (ssthresh) to 0.8 times the
>     FlightSize, with a lower bound of 2 * SMSS applied to the result.
>     As in [RFC5681], the TCP sender also reduces the cwnd value to no
>     more than the new ssthresh value.
> 
> 
> 
> And is further clarified in Section 4.3:
> 
> 
> and in response to an indication of an ECN-signalled congestion:
> 
>       ssthresh = max (FlightSize * beta_{ecn}, 2 * SMSS)
> 
>       and
> 
>       cwnd = ssthresh
> 
> 
> 
> In my reading this text above sets a lower bound of 2 MSS for cwnd which is incorrect and dangerous modification and maybe not intended to modify RFC 3168.

Just to be clear, indeed we absolutely did not intend to change the spec here - however we thought we need to be clearer.

Consider the original text:
***
    The indication of congestion should be treated just as a
    congestion loss in non-ECN-Capable TCP.  That is, the TCP source
    halves the congestion window "cwnd" and reduces the slow start
    threshold "ssthresh”. 
***

Because this simply says “reduces” about ssthresh without specifying in detail how to reduce it, it would seem obvious (due to the first sentence) that this refers to eqn. 4 in RFC 5681. However, if we say “...set the slow start threshold (ssthresh) to 0.8 times the FlightSize”, then we lose the lower bound. This is why we introduced it for ssthresh.

As for cwnd, I find the original text to be contradictory in this respect (unless I’m missing something). “The indication of congestion should be treated just as a congestion loss in non-ECN-Capable TCP. “ => this, to me, means doing the same as specified in RFC 5681 - but in RFC 5681, I couldn’t see any text that allows cwnd to go below 2 MSS upon a “congestion loss” (assuming this is not a timeout - where you’re right, our text wrongly quoted RFC 5681 there, but that’s a separate issue further below).

( Maybe you mean *during*, not after fast recovery? But that would be confusing, as cwnd is also “inflated” in this phase in RFC 5681, which I hope we don’t do in response to ECE. )

However, RFC 3168, in the second sentence, beginning with “That is, “ (indicating that this is just a clarification), states that the TCP sources halves cwnd.


> The lower bound for cwnd is 1 MSS. The lower bound of 2 MSS for cwnd
> 
> 1) is not specified in any (standards track) RFC,
> 2) is against the congestion control principles for avoiding
>   congestion collapse as it effectively disables a full back
>   off that is a crucial requirement for all congestion control
>   algorithms,
> 3) is in conflict with the basic principle set in Sally's ECN papers
>   and in RFC 3168 itself (the original text above from RFC 3168):
> 
>   "The indication of congestion should be treated just as a
>    congestion loss in non-ECN-Capable TCP."
> 
>   When Flightsize (and cwnd) is 2 MSS in non-ECN-Capable TCP and loss
>   signals congestion, cwnd effectively becomes 1 MSS. That is why RFC
>   3168 formulates the text: "the TCP source halves the congestion
>   window "cwnd" …"

See above, regarding this text. We certainly don’t want to change these principles but tried to be correct about what should happen, and found the quoted text here to be contradictory.


>   Note: the majority of the work and text in RFC 3168, including
>   the text in RFC 2482, predates RFC 2581 that first introduced
>   FlightSize. Prior to RFC 2581 all TCP CC text "halved cwnd"
>   which literally meant cwnd=cwnd/2.

I know that, and I thought that this is the reason for the problem with this text (due to updates of the non-ECN CC text).


> Instead of the above text in the draft, the draft should explicitly state that when FlightSize/2 would result in ssthresh below its lower bound of 2 MSS, a TCP sender must not stop decreasing cwnd. And, if more ECN-signalled congestion indications arrive, the TCP sender must continue reducing sending rate further (plus refer to RFC 3168 how to proceed reducing the sending rate, if further ECE signals arrive).

Technically this makes sense to me, but I’m not currently reading RFC 3168 this way.


> Furthermore, the draft should be very specific on how the TCP sender should act when cwnd=ssthresh and ECE arrives. Currently the draft specifies the way how the TCP sender reduces ssthresh (and cwnd) in Congestion Avoidance. According to RFC 5681, when cwnd = ssthresh, TCP is not in Slow Start nor in Congestion Avoidance. Therefore, the draft should be explicit on how cwnd is reduced, when cwnd=ssthresh.

Why? RFC 5681 states “When cwnd and ssthresh are equal, the sender may use either slow start or congestion avoidance.” - so if you're an implementer, this basically tells you that you have to make a decision on your own for the cwnd == sthresh case. We specify how to behave when you’re in congestion avoidance - thus it depends on the decision you’ve taken for this case.


> It is true that there is (at least) one (widely deployed) incorrect implementation of RFC 3168 that incorrectly sets lower bound of 2 MSS for cwnd when congestion is ECN-signaled. This, however, must not be used as any kind of justification for changíng IETF-specified congestion control principles and requirements. Instead, this draft should carefully instruct implementors and those who intend to experiment with the ECN modification set in this draft to ensure the TCP stack used for experiments correctly implements RFC 3168 in this respect. Otherwise, such experimental results are mostly void.

Maybe the person behind this implementation was as confused by the RFC 3168 as I am … clearly we didn’t intend to change the principles.


> In addition, there is an incorrect statement in Sec 4.1:
> 
> According to
> [RFC5681], when a TCP sender detects segment loss using the
> retransmission timer and the given segment has not yet been resent by
> way of the retransmission timer, the value of ssthresh must be set to
> no more of the maximum of half of the FlightSize and 2*SMSS.  The
> same equation is also used during Fast Retransmit/Fast Recovery
> [RFC5681].  As a result, the TCP congestion control only allows one
> BDP of packets in flight.  This is just sufficient to maintain 100%
> utilisation of the bottleneck on the network path.
> 
> 
> What the text says about setting ssthresh is correct. However, if the segment loss is detected using the retransmission timer, the TCP
> congestion control does not allow one BDP of packets in flight (unless BDP=1 MSS). Instead, cwnd is set to 1 MSS, allowing only 1 pkt in flight.
> 
> I suggest removing the text that discusses detecting segment loss using the retransmission timer.

Good catch, thanks!

Cheers,
Michael