Re: [Rift] RIFT Open Source Update Presentation (RIFT WG meeting IETF-108)
Bruno Rijsman <brunorijsman@gmail.com> Mon, 27 July 2020 12:06 UTC
Return-Path: <brunorijsman@gmail.com>
X-Original-To: rift@ietfa.amsl.com
Delivered-To: rift@ietfa.amsl.com
Received: from localhost (localhost [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id F0C803A190D for <rift@ietfa.amsl.com>; Mon, 27 Jul 2020 05:06:51 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -2.097
X-Spam-Level:
X-Spam-Status: No, score=-2.097 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, FREEMAIL_FROM=0.001, HTML_MESSAGE=0.001, SPF_HELO_NONE=0.001, SPF_PASS=-0.001, URIBL_BLOCKED=0.001] autolearn=ham autolearn_force=no
Authentication-Results: ietfa.amsl.com (amavisd-new); dkim=pass (2048-bit key) header.d=gmail.com
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id eq2Q-c53RfHN for <rift@ietfa.amsl.com>; Mon, 27 Jul 2020 05:06:47 -0700 (PDT)
Received: from mail-ej1-x643.google.com (mail-ej1-x643.google.com [IPv6:2a00:1450:4864:20::643]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by ietfa.amsl.com (Postfix) with ESMTPS id 9B5BE3A190C for <rift@ietf.org>; Mon, 27 Jul 2020 05:06:46 -0700 (PDT)
Received: by mail-ej1-x643.google.com with SMTP id f14so946778ejb.2 for <rift@ietf.org>; Mon, 27 Jul 2020 05:06:46 -0700 (PDT)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20161025; h=from:message-id:mime-version:subject:date:in-reply-to:cc:to :references; bh=uUwsXu4jA93DhoKKqgv4AiDZfyP2KjJclC26AHro7FM=; b=S1EJ0+N4A9VxL0wTyyxADlUygrH427Oeh4cPva0nKFO3Q2aimVTWBZmLkEF0oXriv1 hatEq6BgeGME/lKf5iEZGYNY2I4RHTge3NF33rtB7KsB2ArvpTTw6h/tej5zhnfDg50v xzoq+eaxdkwGDpgs7FdQJNpI8nS4f84F1hxPSsWktmjIhPMBHnbvY21mgR68kzjVT9EU X1gHP35hND7VL6ZoLN9wYM6J2/ZcmhOT4/p0aFisSRG73X9mkEMJkNMcHXIeq7/FttD7 oo5kwZBF2FdOUUQ/0NVsmSXb/EXkP4eijfxqPdCyERRkDuLoW0yJFVtFUnYHR5MQTGBl 9rNQ==
X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20161025; h=x-gm-message-state:from:message-id:mime-version:subject:date :in-reply-to:cc:to:references; bh=uUwsXu4jA93DhoKKqgv4AiDZfyP2KjJclC26AHro7FM=; b=YkWz7Lwr2ai/8+YiArWx/+hMNUdJ1+ZywSXCQTegFv9bCybQ3pivUyIOTOeRPu1Mgx 9TnfYfXXS50kJqDZZIGyoeCcBi1gy+51TmXLxFiYhxpnfli6swscuWB81Qezs8/anPwg HLnTt+FfsMq79fUcgD891zuFpjJma/FAmra3Ou3C1NaOr8ZYzhAqHIexNwxyVtlVOAMq xgd3sehf0t2haL7+QnEQlwGdaxuM9XdoXA9pnDeikU2tOQaP42utlL4gEf1HwfO2Z6ou 6+pLBCzIcx19y6OU5f9DxdVBVrYQoke8cDHGNmwsoZ2GuxRU+wngKp0aMfmHjxJj5+Nu s5gg==
X-Gm-Message-State: AOAM533fvtgYRer6BN0kGsU8cJhjI9hb8dMUjP1Cl+3C2flNemfcxWuz Xcaz9hZyBoFitnfRL66SBxQ=
X-Google-Smtp-Source: ABdhPJydiJD7Ycjb3vgJB7w96+K1RnSRPQvNqTN1a7V9lFoKBIpLUDs26rj0k7kgV3YPWxaON1Fw8A==
X-Received: by 2002:a17:907:212b:: with SMTP id qo11mr20477373ejb.452.1595851604902; Mon, 27 Jul 2020 05:06:44 -0700 (PDT)
Received: from [192.168.1.100] (84-31-170-55.cable.dynamic.v4.ziggo.nl. [84.31.170.55]) by smtp.gmail.com with ESMTPSA id v24sm7216354eds.71.2020.07.27.05.06.44 (version=TLS1_2 cipher=ECDHE-ECDSA-AES128-GCM-SHA256 bits=128/128); Mon, 27 Jul 2020 05:06:44 -0700 (PDT)
From: Bruno Rijsman <brunorijsman@gmail.com>
Message-Id: <3FD14A34-5927-490E-A586-C47FA2AD9909@gmail.com>
Content-Type: multipart/alternative; boundary="Apple-Mail=_CDED6E56-7E85-4C3B-B64C-D161152EB294"
Mime-Version: 1.0 (Mac OS X Mail 13.4 \(3608.80.23.2.2\))
Date: Mon, 27 Jul 2020 14:06:43 +0200
In-Reply-To: <CA+b+ERnUxtKvcDVP2KxLiSk3wHR_+1V75oDu=5bpSs2vDpw0dQ@mail.gmail.com>
Cc: rift@ietf.org
To: Robert Raszuk <rraszuk@gmail.com>
References: <A5DE1D42-F878-461A-A64F-84DB8678A46F@gmail.com> <CA+b+ERnUxtKvcDVP2KxLiSk3wHR_+1V75oDu=5bpSs2vDpw0dQ@mail.gmail.com>
X-Mailer: Apple Mail (2.3608.80.23.2.2)
Archived-At: <https://mailarchive.ietf.org/arch/msg/rift/hsfvfqwdxPIJ__d5ZxKeI2gKes0>
Subject: Re: [Rift] RIFT Open Source Update Presentation (RIFT WG meeting IETF-108)
X-BeenThere: rift@ietf.org
X-Mailman-Version: 2.1.29
Precedence: list
List-Id: Discussion of Routing in Fat Trees <rift.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/rift>, <mailto:rift-request@ietf.org?subject=unsubscribe>
List-Archive: <https://mailarchive.ietf.org/arch/browse/rift/>
List-Post: <mailto:rift@ietf.org>
List-Help: <mailto:rift-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/rift>, <mailto:rift-request@ietf.org?subject=subscribe>
X-List-Received-Date: Mon, 27 Jul 2020 12:06:53 -0000
Hi Robert, Thanks for the observations and questions! See answers below. — Bruno > On Jul 25, 2020, at 8:56 PM, Robert Raszuk <rraszuk@gmail.com> wrote: > > Hi Bruno, > > This is excellent ! I wish everyone would prepare such video ahead of time like you did. > > One observation and few questions ... > > * First let me observe that concept of negative routing in the control plane is not new. As example - the proposals on that have been written many years ago for BGP. Perhaps it is unfortunate that vendor did not choose to pursue it in public at that time :) But it did pursue patents and implementation for negative routing over low power and lossy networks. > I would be very interested in the negative routing proposal that was written for BGP (or for any other protocol). Do you have any public references? Did it also include an automatic mechanism for triggering negative disaggregation, or did it require manual configuration of the negative routes to be advertised? (Since you mention the vendor did not pursue this in public, I would understand if you don't have a reference.) Either way, I certainly did not intend to imply that negative disaggregation has never been done before. In technology, as in art, almost everything has been done before. It's just that the publicly available implementations of BGP, ISIS, OSPF, etc. currently don't include the concept of negative disaggregation, or at least not as far as I know. Even within the narrow scope of RIFT, I don't claim to have contributed any original thoughts on the topic of negative disaggregation -- the vast majority of the credit for that goes to Pascal Thubert (it is perhaps no coincidence that he works on low power and lossy networks for a major network equipment vendor, perhaps the same one that you allude to). > * In your talk you describe the behaviour in rest of the network in regards to negative routing. But how will leaf-1 in pod-1 receive negative routes if the real trigger for them is based on control plane messages traveling over horizontal links between super spines ? Leaf-1-1 in pod-1 does not rely on negative disaggregation to avoid plane-1. Instead it relies on the fact that RIFT originates a south-bound default route at the top-of-fabric nodes, and then lower nodes in the hierarchy only propagate the south-bound default route if they have received a default route advertisement from at least one parent. This is a little bit of a simplification, see section 4.2.3.8in version 12 of the RIFT draft for details, and there is also an interesting interaction with ZTP. See APPENDIX below for details. > > * How do you stabilize/control control plane and data plane churn during negative routing being triggered by link flapping (say one down and one flapping in your example between pod 1 and plane 1) and no native link dampening in place ? > TIEs (= RIFT Link State Packets) for negative disaggregate prefixes are subject to the same link-flap considerations that apply to normal node TIEs or normal positive prefix TIE. Hence, all the usual mechanisms that protect link-state protocols from excessive link flapping apply, including but not limited to: * Limiting the rate at which TIEs are flooded over adjacencies. In RIFT there are transmit queues for TIEs, and a distinction is made between fast transmits (for the initial transmission) and slow re-transmits (for retransmits). * Limiting the rate at which SPF runs are conducted. * Various flow-control mechanisms (including sequence numbers and the `you_are_sending_too_quickly` field in LIEs) to detect that RIFT packets are being dropped and adjusting the flooding rates accordingly. (I have not yet implemented this in my code -- see also PS below). > * Is the load balancing spread based on the configured link bandwidth for the interfaces ? If so it does not account for brownouts where fabric in any network element itself or linecard partially fails and while say BFD or RIFT still goes through the actual bandwidth for the data plane is severely reduced. Any plans for this little enhancement ? :) > In my code, the load balancing is indeed based on the configured interface speed. That said, it is an implementation choice, and there is nothing in the RIFT draft that prevents an implementation from doing something more sophisticated, for example using bandwidth based on actual telemetry measurements. I currently have no plans to do that -- see also PS below. > Thank you, > R. > > PS. The FSM performance is a fantastic addition. I wish all code would have such CLI :) But while we are at this is there some analysis how would python-rift performance compare with say c-rift when dealing with say 1M to 10M routes fabric ? Case of no overlay and flat routing of /32s and /128s all over the cluster. > The RIFT-Python code was originally started with the following goals in mind: * Check whether it was possible to create an interoperability implementation of RIFT based purely on the text in the RIFT specification, and where necessary, improve the quality of the RIFT specification. * Be a reference implementation of RIFT. It was never a goal of the RIFT-Python project to support 1000s of routes or millions of routes. The code is optimized for clarity and readability. It is not optimized for performance: it is not multi-threaded and Python is not a logical choice for performance. That said, I expect that RIFT-Python would make a fine host-based RIFT router since the RIFT performance demand on leaf routers is very modest, even in very large topologies. APPENDIX Let's go through the concrete example shown on slide 12 in the presentation deck. Before the links are broken, the sitation is as follows: Node super-1-1 originates a south-bound default route because it is top-of-fabric: super-1-1> show node Node: +-------------------------------------------+-------------------------------------------------------------+ | Name | super-1-1 | +-------------------------------------------+-------------------------------------------------------------+ | Originate IPv4 Default Route | True | | Reason for Originating IPv4 Default Route | All other nodes at my level have no north-bound adjacencies | | Originate IPv6 Default Route | True | | Reason for Originating IPv6 Default Route | All other nodes at my level have no north-bound adjacencies | +-------------------------------------------+-------------------------------------------------------------+ Node spine-1-1 originates a south-bound default route (as do spine-1-2 and spine-1-3) because it received a default advertisement from the superspine and computed a default route to install in its RIB: spine-1-1> show node Node: +-------------------------------------------+----------------------------------------------+ | Name | spine-1-1 | +-------------------------------------------+----------------------------------------------+ | Originate IPv4 Default Route | True | | Reason for Originating IPv4 Default Route | This node has north-bound IPv4 default route | | Originate IPv6 Default Route | True | | Reason for Originating IPv6 Default Route | This node has north-bound IPv6 default route | +-------------------------------------------+----------------------------------------------+ Hence, leaf-1-1 has north-bound default routes that ECMP the traffic over all three planes (for the sake of space I only show IPv4 routes): leaf-1-1> show route IPv4 Routes: +-----------+-----------+----------+-----------------+----------+----------+ | Prefix | Owner | Next-hop | Next-hop | Next-hop | Next-hop | | | | Type | Interface | Address | Weight | +-----------+-----------+----------+-----------------+----------+----------+ | 0.0.0.0/0 | North SPF | Positive | veth-1001a-101a | 99.0.2.2 | 33 | | | | Positive | veth-1001b-102a | 99.0.4.2 | 33 | | | | Positive | veth-1001c-103a | 99.0.6.2 | 33 | +-----------+-----------+----------+-----------------+----------+----------+ After we break links spine-1-1 <-> super-1-1 and spine-1-1 <-> super-1-2, we would expect the following to happen: 1. Super-1-1 still advertises a south-bound default route. 2. But spine-1-1 doesn't receives the default advertisement from any superspine (since all links are broken) and hence does not propage the default advertisement: 3. Hence, leaf-1-1 now has a north-bound default route that only ECMPs over plane-2 and plane-3, but avoids plane-1: leaf-1-1> show route IPv4 Routes: +-----------+-----------+----------+-----------------+----------+----------+ | Prefix | Owner | Next-hop | Next-hop | Next-hop | Next-hop | | | | Type | Interface | Address | Weight | +-----------+-----------+----------+-----------------+----------+----------+ | 0.0.0.0/0 | North SPF | Positive | veth-1001b-102a | 99.0.4.2 | 50 | | | | Positive | veth-1001c-103a | 99.0.6.2 | 50 | +-----------+-----------+----------+-----------------+----------+----------+ However, before we even get there something else quite surprising happens: 1. When spine-1-1 loses both north-bound adjacencies, it is no longer able to auto-derive it's level using ZTP. 2. This causes spine-1-1 to shut down all of its adjacencies including the south-bound ones to the leaves. 3. This has the same (good) result: leaf-1-1 now has a north-bound default route that only ECMPs over plane-2 and plane-3, but avoids plane-1. This was discussed on the RIFT mailing list on 23-April-2020 (see https://mailarchive.ietf.org/arch/msg/rift/zGLnispMv630LWq0ALRgGs7dLJI/) > > On Sat, 25 Jul 2020 at 18:44, Bruno Rijsman <brunorijsman@gmail.com <mailto:brunorijsman@gmail.com>> wrote: > Because IETF-108 is a pure virtual meeting, I went through the trouble of pre-recording my "RIFT Open Source Update" presentation scheduled for the RIFT working group meeting on Wednesday. > > Enjoy the "Sneak preview" ;-) > > Video: https://www.youtube.com/watch?v=qhiiTDPuku0 <https://www.youtube.com/watch?v=qhiiTDPuku0> > > Slides: https://bit.ly/rift-open-source-update-ietf-108 <https://bit.ly/rift-open-source-update-ietf-108> > > — Bruno Rijsman > _______________________________________________ > RIFT mailing list > RIFT@ietf.org <mailto:RIFT@ietf.org> > https://www.ietf.org/mailman/listinfo/rift <https://www.ietf.org/mailman/listinfo/rift>
- [Rift] RIFT Open Source Update Presentation (RIFT… Bruno Rijsman
- Re: [Rift] RIFT Open Source Update Presentation (… Robert Raszuk
- Re: [Rift] RIFT Open Source Update Presentation (… Bruno Rijsman