[tsvwg] Re: Request to review diffserv spec: draft-ietf-tsvwg-nqb

Vasilenko Eduard <vasilenko.eduard@huawei.com> Fri, 31 May 2024 11:07 UTC

Return-Path: <vasilenko.eduard@huawei.com>
X-Original-To: tsvwg@ietfa.amsl.com
Delivered-To: tsvwg@ietfa.amsl.com
Received: from localhost (localhost [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 2246CC14F714 for <tsvwg@ietfa.amsl.com>; Fri, 31 May 2024 04:07:33 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -6.395
X-Spam-Level:
X-Spam-Status: No, score=-6.395 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, GB_ABOUTYOU=0.5, RCVD_IN_DNSWL_HI=-5, RCVD_IN_MSPIKE_H3=0.001, RCVD_IN_MSPIKE_WL=0.001, RCVD_IN_ZEN_BLOCKED_OPENDNS=0.001, SPF_HELO_NONE=0.001, SPF_PASS=-0.001, URIBL_DBL_BLOCKED_OPENDNS=0.001, URIBL_ZEN_BLOCKED_OPENDNS=0.001] autolearn=unavailable autolearn_force=no
Received: from mail.ietf.org ([50.223.129.194]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id vtMUSZfC-xJN for <tsvwg@ietfa.amsl.com>; Fri, 31 May 2024 04:07:29 -0700 (PDT)
Received: from frasgout.his.huawei.com (frasgout.his.huawei.com [185.176.79.56]) (using TLSv1.2 with cipher AECDH-AES256-SHA (256/256 bits)) (No client certificate requested) by ietfa.amsl.com (Postfix) with ESMTPS id 564E4C15108C for <tsvwg@ietf.org>; Fri, 31 May 2024 04:07:29 -0700 (PDT)
Received: from mail.maildlp.com (unknown [172.18.186.31]) by frasgout.his.huawei.com (SkyGuard) with ESMTP id 4VrKwV0RkMz6K6Fn; Fri, 31 May 2024 19:03:06 +0800 (CST)
Received: from mscpeml100004.china.huawei.com (unknown [7.188.51.133]) by mail.maildlp.com (Postfix) with ESMTPS id D26AA140A08; Fri, 31 May 2024 19:07:25 +0800 (CST)
Received: from mscpeml500004.china.huawei.com (7.188.26.250) by mscpeml100004.china.huawei.com (7.188.51.133) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.2.1258.34; Fri, 31 May 2024 14:07:25 +0300
Received: from mscpeml500004.china.huawei.com ([7.188.26.250]) by mscpeml500004.china.huawei.com ([7.188.26.250]) with mapi id 15.02.1258.034; Fri, 31 May 2024 14:07:25 +0300
From: Vasilenko Eduard <vasilenko.eduard@huawei.com>
To: Sebastian Moeller <moeller0=40gmx.de@dmarc.ietf.org>
Thread-Topic: [tsvwg] Request to review diffserv spec: draft-ietf-tsvwg-nqb
Thread-Index: AQHassUh51dXLRlIeUG0gVItsRaLwrGw62WA///kiYCAAF4HgA==
Date: Fri, 31 May 2024 11:07:25 +0000
Message-ID: <66a648b31faf4135ac7794a340d42ec2@huawei.com>
References: <171619088820.9700.17122047615729502291@ietfa.amsl.com> <65a001f0-386e-4d3a-b41a-1ae75a195aca@erg.abdn.ac.uk> <075c9da4-df83-4286-85f3-d36553fac3ba@gmail.com> <aa71804ee5194509914e74e717766247@huawei.com> <HrBN5VBqLPXzuh9BCMrR8cWB1fKo5aGo3Axw3nT9XQDftHqNXdtWFvnUBt0N8N3YbC_Mj77KPzk_UbGvlPbq3pcIdC1m4tBg2kmf9BPk6uM=@ealdwulf.org.uk> <20e06b97714b412681a2d444525c928b@huawei.com> <950432d946054b8da5af4c4b37caa5b7@huawei.com> <CEDCA860-5A21-4E5F-BB9E-EA45F6233388@CableLabs.com> <6F643FEF-8664-4B79-B71D-0A70A6DF4723@gmx.de> <0e7cfe1a05cb4a4c9a7d89b7e8a69743@huawei.com> <D0116F64-8D66-4915-90A8-FB90B52AF058@gmx.de>
In-Reply-To: <D0116F64-8D66-4915-90A8-FB90B52AF058@gmx.de>
Accept-Language: en-US
Content-Language: en-US
X-MS-Has-Attach:
X-MS-TNEF-Correlator:
x-originating-ip: [10.199.59.222]
Content-Type: text/plain; charset="utf-8"
Content-Transfer-Encoding: base64
MIME-Version: 1.0
Message-ID-Hash: ALZQW5AKKRBNIV4YHZNBIOYZZR7NGMOL
X-Message-ID-Hash: ALZQW5AKKRBNIV4YHZNBIOYZZR7NGMOL
X-MailFrom: vasilenko.eduard@huawei.com
X-Mailman-Rule-Misses: dmarc-mitigation; no-senders; approved; emergency; loop; banned-address; member-moderation; header-match-tsvwg.ietf.org-0; nonmember-moderation; administrivia; implicit-dest; max-recipients; max-size; news-moderation; no-subject; digests; suspicious-header
CC: "tsvwg@ietf.org" <tsvwg@ietf.org>
X-Mailman-Version: 3.3.9rc4
Precedence: list
Subject: [tsvwg] Re: Request to review diffserv spec: draft-ietf-tsvwg-nqb
List-Id: Transport Area Working Group <tsvwg.ietf.org>
Archived-At: <https://mailarchive.ietf.org/arch/msg/tsvwg/Aow25pHwqh91kHUY-PnWGfSd5To>
List-Archive: <https://mailarchive.ietf.org/arch/browse/tsvwg>
List-Help: <mailto:tsvwg-request@ietf.org?subject=help>
List-Owner: <mailto:tsvwg-owner@ietf.org>
List-Post: <mailto:tsvwg@ietf.org>
List-Subscribe: <mailto:tsvwg-join@ietf.org>
List-Unsubscribe: <mailto:tsvwg-leave@ietf.org>

Hi Sebastian,
I do not understand why you are mixing L4S and NQB in one solution. It does not make sense.

L4S is signaling a special type of service by ECN(1). It is enough to go to the L queue. L queue on L4S is really dynamic - it could consume up to 100% of the bottleneck.

NQB is signaling a special type of service by DSCP45. It is enough to go to a special queue with absolute forwarding priority. Unfortunately, this queue is limited to 5% of the bottleneck. Security problem comes with this fixed queue performance.

I could not imagine the monstrous solution that would do both. Does not matter that it was produced by the same SDO.
IMHO: these are competitive solutions. Both would not work but for very different reasons.
Eduard
-----Original Message-----
From: Sebastian Moeller <moeller0=40gmx.de@dmarc.ietf.org> 
Sent: Friday, May 31, 2024 11:24
To: Vasilenko Eduard <vasilenko.eduard@huawei.com>
Cc: tsvwg@ietf.org
Subject: Re: [tsvwg] Request to review diffserv spec: draft-ietf-tsvwg-nqb

Hi Ed,

I trimmed the CC list a bit, but left the list included.


> On 31. May 2024, at 09:24, Vasilenko Eduard <vasilenko.eduard=40huawei.com@dmarc.ietf.org> wrote:
> 
> Hi Sebastian,
> I believe that you have given the wrong example for an attack scenario.

[SM] Well possible.

> You did mention DualQ and L4S that have a fully dynamic bandwidth split between queues, if a high-priority queue is more utilized - it could consume up to 100% of bandwidth. Under such a scenario, a hacker would still need to supply 100% of traffic, no additional gain for him.

[SM] That is not the best way of looking at that. At any given moment in time a typical link is either 100% used or 0% so we need to look at the duty cycle and perform sone averaging to figure out required saturation capacity. At the network edge we typically have a noticeable step-down in capacity between the 'backbone' side and the 'customer' side. That is the transfer time for a given volume is considerably shorter on the backbone side than on the customer side.
With that out of the way, IMHO the issue is quite clear, as long as I can push in say 2 * queue-depth worth of bytes into the L-queue in less than 2 * queue-depth worth of time I can drive the queue into action (marking/dropping/re-queueing to some other queue) independent of what is going on in the C-queue. To disrupt that queue I now only need to repeatedly burst 2*queue-depth worth of bytes into that queue with a rate tailored to the traffic I want to attack. The C-queue traffic would not be affected all that much by this, by virtue of a deeper queue that can absorb the burst shock better (with the L-queue having conditional priority over the C-queue even the C-queue traffic will notice these bursts, but will not be affected as much).

*) Picked just as an example of a volume clearly sufficient to fill the shallow queue past its 'action' threshold.


> NQB proposed a fixed split (5% of the interface speed) for a low-latency queue, everything above would be dropped. Then hackers would need to push 20x less traffic to block all applications that use DSCP45.
> Legacy applications (DSCP0) would not be affected in NQB.

[SM] Sure that makes things even easier... not sure though that this is actually a tested deployment and not just a theoretical consideration to decouple NQB from L4S... in the context of low latency docsis NQB will likely be implemented as part of the same AQM/scheduler intenden for ECT(1)/L4S traffic, so I would expect that the 'NQB as part of L4S' likely will see more deployment than the 5% policer/shaper described in the paper. (But I notice that the proposal in the draft mentions 10ms burst tolerance, still something that can be aimed for without a full fledged volumetric attack.)

> Hence, NQB could create a motto: "If you do not want to be effectively 
> DoSed - do not mark your packets with DSCP45". (As I understand - it 
> is the situation now)
> 
> I am still not very concerned about NQB on access (despite NQB being a very big improvement for hackers) because I see little motivation for DoS on access.
> Except for gamers, but If gamers would block each other - the world would become better, right? They would have to do something more useful for mankind.

[SM] Well, the issue here is that e.g. low-latency docsis explicitly targets gamers https://www.cablelabs.com/technologies/low-latency-docsis starts with:
"Low Latency Is Here, Not a Second Too Late From gamers to medical patients to everyday internet users, customers are seeking real-time control over their applications-without choppiness, freezing or other latency issues that won't be improved by simply adding more bandwidth. CableLabs' Low Latency DOCSIS (LLD) technology is a new approach that targets a reduction in round-trip latency in the DOCSIS network to sub-5 ms at the 99th percentile for non-queue building applications. This enhancement to cable broadband will ensure that web pages load faster, video calling is smooth and online gaming is highly responsive."

and the white paper (https://community.cablelabs.com/wiki/plugins/servlet/cablelabs/alfresco/download?id=0affb960-25f4-44f9-b747-74ad8555fd06) mentions gaming as a use-case a lot.

And ComCast's L4S tests (https://github.com/jlivingood/IETF-L4S-Deployment/blob/main/Week1.md) did spend considerable time on L4S/NQB in the context of gaming (one of 5 weeks seemed to have focused mostly on gaming).

And IMHO from an ISP perspective that makes a ton of sense, gaming as an industry overtook Hollywood in revenue some years ago, and gamers are willing to spend considerable amounts of money on their hobby/profession. Whether you or I consider that to be useful for humanity or not, that is a relatively well defined market segment that both L4S and NQB seemed to have aimed at.


> In general, I am voting for a mechanism blocking gamers. For this to happen:
> - NQB should be implemented only on BNG, CMTS, UPF, WiFi gateway 
> typically used in hotels (in general NQB does not make sense for any 
> other place in networking except access),

[SM] Queue management in general, IMHO makes sense where ever noticeable queues build up independent of capacity/rate. It might be infeasible/undesirable for some locations but that is simply a trade-off between expected gain and required cost. For the case of NQG specifically I agree that the most likely bottlenecks will be at the edge and for a given path queue management only needs to happen at the bottleneck to be effective. (Sidenote, for many access links now the home WiFi is becoming the bottleneck so queue management needs to move into the APs and due to WiFis often decentralised nature also into the stations.)

> - gaming applications have to use DSCP45.
> Then gamers would start to attack each other.
> Unfortunately, Gaming applications would disable DSCP45 after this.

[SM] Well possible, that for a game maker sticking to EF might be the better choice, but that depends on whether the gains from a shallow queue overcome the additional risk of a shallow queue

> Eduard
> -----Original Message-----
> From: Sebastian Moeller <moeller0=40gmx.de@dmarc.ietf.org>
> Sent: Thursday, May 30, 2024 22:10
> To: Greg White <g.white@CableLabs.com>
> Cc: Vasilenko Eduard <vasilenko.eduard@huawei.com>; tsvwg@ietf.org; 
> Brian E Carpenter <brian.e.carpenter@gmail.com>; Gorry Fairhurst 
> <gorry@erg.abdn.ac.uk>; Black, David <David.Black@dell.com>
> Subject: Re: [tsvwg] Request to review diffserv spec: 
> draft-ietf-tsvwg-nqb
> 
> 
> 
>> On 30. May 2024, at 20:39, Greg White <g.white=40CableLabs.com@dmarc.ietf.org> wrote:
>> 
>> Hi Eduard,
>> 
>> Thanks for your review and comments.
>> 
>> To your comment that traffic protection and even the NQB PHB itself is not appropriate for "big Internet resources" (which I interpret to be core network and interconnect switches), the draft agrees with you:
>> 
>> https://www.ietf.org/archive/id/draft-ietf-tsvwg-nqb-23.html#section-
>> 1
>> -5 "This PHB is primarily applicable for high-speed broadband access 
>> network links, where there is minimal aggregation of traffic, and deep buffers are common."
>> 
>> https://www.ietf.org/archive/id/draft-ietf-tsvwg-nqb-23.html#section-
>> 4
>> .2-3 "In nodes that do not typically experience congestion (for 
>> example, many backbone and core network switches), forwarding packets with the NQB DSCP using the Default treatment might be sufficient to preserve loss/latency/jitter performance for NQB traffic."
>> 
>> https://www.ietf.org/archive/id/draft-ietf-tsvwg-nqb-23.html#section-
>> 5
>> .2-10 "There are some situations where traffic protection is 
>> potentially not necessary. .... Another example could be highly aggregated links (links designed to carry a large number of simultaneous microflows), where individual microflow burstiness is averaged out and thus is unlikely to cause much actual delay."
>> 
>> As Alex pointed out, one algorithm for traffic protection on access network segments (the docsis queue protection algo) keeps flow state only for non-compliant flows.  Certainly, it is still possible to exhaust this flow state, as discussed in https://www.ietf.org/archive/id/draft-briscoe-docsis-q-protection-07.html#name-resource-exhaustion-attacks  But, there are also counter-measures for this (which have been implemented in proprietary extensions).  And, nothing is perfect in the world of DoS/DDoS prevention for access network segments anyway. For example, single queue FIFOs can easily be attacked (as you mentioned), and FQ-CoDel can be specifically attacked by flooding traffic with incrementing port numbers.
> 
> [SM] "flooding traffic" is the operational word here, that is you need a volumetric attack that exhausts the capacity over timeframes above the queue depth... for the L-queue far less volume is necessary as you can attack with bursts that are enough to drive the L-queue beyond its mark/drop threshold... (incrementing port numbers can likely also be employed overwhelm queue-protections blame assignment scheme). 
> Current Linux DualQ is a great example here as the default queue depth is 1 millisecond after which the step AQM will mark all ECT(1)/CE packets with sojourn times above 1 ms and drop all packets with Not-ECT (or ECT(0)?), so a burst worth > 1ms of capacity will affect L-queue traffic and cause dropping of NQB packets (without ECT(1) or CE).
> 
> And even with out the L4S scheduler in play but something closer to the sketch given in the NQB draft ("A traffic shaping function SHOULD allow approximately 10 ms of burst tolerance, and no more than 50 ms of buffering".) the problem would persist. As a direct consequence of a shallow*deep queue pair is that the shallower queue is more sensitive for short term overload than the deeper queue ... and the corollary is that attack traffic targeting the shallow queue now does not need to overwhelm the full link capacity, but it is sufficient to overwhelm the shallow queue. And with an unpoliced admission to the shallow queue by DSCP dec 45 targeting the shallow queue is not a challenge either.
> 
> 
>> I'm curious about your assertion that: "NQB would help them to DoS much more effectively (20x less traffic)".  Could you provide some evidence to back this up?
> 
> [SM] Good question. Do you have evidence to counter my claim above?
> 
> Regards
> Sebastian
> 
>> 
>> -Greg
>> 
>> 
>> On 5/27/24, 11:20 PM, "Vasilenko Eduard" <vasilenko.eduard=40huawei.com@dmarc.ietf.org <mailto:40huawei.com@dmarc.ietf.org>> wrote:
>> 
>> 
>> Hi all,
>> I would like to state it again:
>> NQB currently creates an impression that this security issue is solvable later.
>> My point: it is not! Only by full replacement of NQB by flow-based AQM (like FQ-Codel).
>> 
>> 
>> But the primary purpose of my message is to state that Sebastian Moeller almost persuaded me off-line that it is a critical issue for subscribers too.
>> Games DoS each other, NQB would help them to DoS much more effectively (20x less traffic).
>> I am just a little hesitant to accept that gaming is something important.
>> Eduard
>> -----Original Message-----
>> From: Vasilenko Eduard
>> Sent: Monday, May 27, 2024 13:41
>> To: 'Alex Burr' <alex.burr@ealdwulf.org.uk 
>> <mailto:alex.burr@ealdwulf.org.uk>>
>> Cc: Brian E Carpenter <brian.e.carpenter@gmail.com 
>> <mailto:brian.e.carpenter@gmail.com>>; tsvwg@ietf.org 
>> <mailto:tsvwg@ietf.org>; Gorry Fairhurst <gorry@erg.abdn.ac.uk 
>> <mailto:gorry@erg.abdn.ac.uk>>; Black, David 
>> <David.Black@dell.com<mailto:David.Black@dell.com>>
>> Subject: RE: [tsvwg] Re: Request to review diffserv spec: 
>> draft-ietf-tsvwg-nqb
>> 
>> 
>> Academia has a *huge* number of research with the keywords "elephant" and "mice". Just google if interested. To get to the point immediately add "monitoring" to the search.
>> They mostly have small tables for loosely aggregated "elephants" and heavily aggregated 'mice' with the mechanism to promote between tables (sometimes very advanced, Bloom filter or even more complex).
>> In all cases it is something small (like 10^4 flows overall on the link) and for a small proportion of "elephants" (single digit X%).
>> Anyway, even after this it consumes considerable resources, the implementation is shown for P4 best case or Xeon typically.
>> Flow-based solutions could be optimized, but it has limits. I am pessimistic about having it scalable even for 100GE links.
>> 
>> 
>> What is good for "monitoring" is not good for AQM.
>> 
>> 
>> Of course, it is possible to read draft-briscoe-docsis-q-protection. Just I doubt that after so many academic papers it is possible to break through this problem.
>> Eduard
>> -----Original Message-----
>> From: Alex Burr <alex.burr@ealdwulf.org.uk 
>> <mailto:alex.burr@ealdwulf.org.uk>>
>> Sent: Monday, May 27, 2024 13:24
>> To: Vasilenko Eduard <vasilenko.eduard@huawei.com 
>> <mailto:vasilenko.eduard@huawei.com>>
>> Cc: Brian E Carpenter <brian.e.carpenter@gmail.com 
>> <mailto:brian.e.carpenter@gmail.com>>; tsvwg@ietf.org 
>> <mailto:tsvwg@ietf.org>; Gorry Fairhurst <gorry@erg.abdn.ac.uk 
>> <mailto:gorry@erg.abdn.ac.uk>>; Black, David 
>> <David.Black@dell.com<mailto:David.Black@dell.com>>
>> Subject: Re: [tsvwg] Re: Request to review diffserv spec: 
>> draft-ietf-tsvwg-nqb
>> 
>> 
>> 
>> 
>> Hi Vasilenko, tsvwg
>> 
>> 
>> See [AB] inline
>> 
>> 
>> On Monday, 27 May 2024 at 07:26, Vasilenko Eduard <vasilenko.eduard=40huawei.com@dmarc.ietf.org <mailto:40huawei.com@dmarc.ietf.org>> wrote:
>> 
>> 
>>> 
>>> 
>>> Hi all,
>>> Brian has attracted my attention to the draft.
>>> 
>>> I believe that the document has a small contradiction to itself:
>>> It is rightly stated that per-flow AQM is not scalable, the separate queue is proposed instead.
>>> At the same time, the protection proposed against malicious actors is proposed per-flow (I could not propose a better one too).
>>> Hence, we could implement something per-flow (like FQ-Codel) in the first place - we do not need this mechanism, it looks redundant.
>> 
>> 
>> [AB] the draft refers to an example mechanism, draft-briscoe-docsis-q-protection. 
>> I'm not an author of either draft, and I don't want to claim anything on their behalf. But one thing that's of interest is that that mechanism, if I've understood it correctly, scales not in the number of flows overall but with the number of non-compliant flows. It seems plausible that only a small, fixed number of resources for non-compliant flows would provide a sufficient incentive for many scenarios.
>> Eg, if a developer classifies their flows as NQB without reading the docs properly, and finds that their application is punished, they are most likely to revert the classification.
>> I'm not sure that draft-briscoe-docsis-q-protection recommends a sufficient punishment, but that can be discussed. There is also some discussion of its behavior under DoS.
>> 
>> 
>> Alex
> 
>