[Idr] Re: I-D Action: draft-vroonen-idr-bgp-bestpath-nh-selection-01.txt

Robert Raszuk <robert@raszuk.net> Thu, 05 March 2026 14:16 UTC

Return-Path: <robert@raszuk.net>
X-Original-To: idr@mail2.ietf.org
Delivered-To: idr@mail2.ietf.org
Received: from localhost (localhost [127.0.0.1]) by mail2.ietf.org (Postfix) with ESMTP id A5CA2C4F4523 for <idr@mail2.ietf.org>; Thu, 5 Mar 2026 06:16:56 -0800 (PST)
X-Virus-Scanned: amavisd-new at ietf.org
X-Spam-Flag: NO
X-Spam-Score: -2.099
X-Spam-Level:
X-Spam-Status: No, score=-2.099 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, HTML_MESSAGE=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_HELO_NONE=0.001, SPF_PASS=-0.001] autolearn=ham autolearn_force=no
Authentication-Results: mail2.ietf.org (amavisd-new); dkim=pass (2048-bit key) header.d=raszuk.net
Received: from mail2.ietf.org ([166.84.6.31]) by localhost (mail2.ietf.org [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id dygWBAQz6oHs for <idr@mail2.ietf.org>; Thu, 5 Mar 2026 06:16:54 -0800 (PST)
Received: from mail-ed1-x531.google.com (mail-ed1-x531.google.com [IPv6:2a00:1450:4864:20::531]) (using TLSv1.3 with cipher TLS_AES_128_GCM_SHA256 (128/128 bits) key-exchange X25519 server-signature ECDSA (P-256) server-digest SHA256) (No client certificate requested) by mail2.ietf.org (Postfix) with ESMTPS id B1DF5C4F426E for <idr@ietf.org>; Thu, 5 Mar 2026 06:16:40 -0800 (PST)
Received: by mail-ed1-x531.google.com with SMTP id 4fb4d7f45d1cf-660fc3f30c1so2647400a12.1 for <idr@ietf.org>; Thu, 05 Mar 2026 06:16:40 -0800 (PST)
ARC-Seal: i=1; a=rsa-sha256; t=1772720200; cv=none; d=google.com; s=arc-20240605; b=bzDRDHMOGz4TpwzavT3dRDU0d9urcrOKfWvmdk2/UD0WZzl4INFqIAf+A2O05/8CTB 0lfMNwuiQSgfJAorQE63GQAxKGT8UAY0ZgGB0ykvsGkKVLODMSHBBwrKOoNftdh+oOI7 RyMx+h9/e69I18HbKF5ZRYUliDGomiF4gHAZ23Ao7n45HBOiCVr05nuJNt5Ldbx/2HUN OKqQa+8LOxgY6beMOWQ7au9gE8+hqEcB0MGigKKcVTw1zrptTFBkc+Ds9pO/oqkQxdkj sQZSka5qKBWpJaiKC35B4cIRaTqoRufg/1Em0QNLnBDwKHvEEruuBiTdZi4DAThVz6to sCRA==
ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=arc-20240605; h=cc:to:subject:message-id:date:from:in-reply-to:references :mime-version:dkim-signature; bh=XzUmmqF68xibg1QSyaJJWpRDxwU16pgGoVREKFlnWRs=; fh=v+Q3yVDYdJWWWkOb7pVJhXIln9TqwuemWzHivLpV/dY=; b=cZJBWRZi/ATnXqnNmDMVZf19GIXCkXIZLS98HhfdEJ25p5l8W4Pg5v6PHgNIvexVxL CmULeEfURDKh5H231JPY7+O8r3CccWXW08gbcpfUcVv3uTfwsWRDPOB3BJ5foXkZMZG/ POrPbiIhZeuOXJua4ZX91FNsG6obXM7zLdbKvuvDIKw6VblzzzyY1LfR9C4fKWMi9vyR zaYNzDxXYabB8XUe1y5wo0foyxrFw7niEOVrQfDIXeT0RYNnjPvQcAOGlHhWI30hoetN g4+M9M8luTKppN9kg83YIdaEZ9DSGfSdfC3xeCiofyJ7ywoVqcZQJ2Nd6R1BOZipuK1H dmMA==; darn=ietf.org
ARC-Authentication-Results: i=1; mx.google.com; arc=none
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=raszuk.net; s=google; t=1772720200; x=1773325000; darn=ietf.org; h=cc:to:subject:message-id:date:from:in-reply-to:references :mime-version:from:to:cc:subject:date:message-id:reply-to; bh=XzUmmqF68xibg1QSyaJJWpRDxwU16pgGoVREKFlnWRs=; b=AVTjenchVtLn0vPiE4NWvM/IE3MzX08s5n8N1o+r+eOSuFHKT38iMau9kYmMdYT8YG iU5o4GxuClZ2RJu7UZzpw3Oyr7JbJhL4zYnL8h+sQMYgXTIhkgPf4v+dzvo4V+vyGEjc jhTq5ZpGo96SNlDMsZCZSUqDP8SmWwHH24WmfiTdVAcHiQiD57Ur+qhXIYDaIqzN71hP 8jV+G4Ki6SNudAInmG0LpQq68NB8Pvbz4CrEClmfpNEu75iPb7cyboo+guzELDSZmI8v 3p7JBTQLzId2cAEVHf1YPWFoFdPIQycz8m+oxeZPYWVeAMyczThirfDlwnoLbGTZcYad ffeg==
X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1772720200; x=1773325000; h=cc:to:subject:message-id:date:from:in-reply-to:references :mime-version:x-gm-gg:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to; bh=XzUmmqF68xibg1QSyaJJWpRDxwU16pgGoVREKFlnWRs=; b=sVwCw6asqUU1pwOGvQLqafhKMgKKYk7o5VnLxXD66nXf5vcow9S4a7QCY/FRKc9lV8 7J/5QRadHmp103e2ytRnrKeZO0z6IDPrtx4mfrQiYEYDTch8VyQoDn5T0zU97UnAk8YE BqD705ekDPZfjpv29AKZLwTRbzxnbFo1RfXFmQb/F9fdklgnZ1P7b5udjixULirtOK0r bM3Tp7ZB51ljaTWiLygWIln3z9+fvme0J4UoetPtUNXHDF0znXj1x84Jy72eKVY9Bin7 BhUQc9iDdwLANRNyaCJJB2MYIxrTUchl30ZDg6hj91tS1rGvk82RYJuMVSTOwLCmY4+k AEog==
X-Forwarded-Encrypted: i=1; AJvYcCU0ETWN5wVUwdXt6r8Ucfz7KpOa3QbsPPPRGEerBODmcFu4IuZ8R7yam1FMiohzjdi8mgE=@ietf.org
X-Gm-Message-State: AOJu0Yx0RKrhzO9BsXwPMjlW3FMrnVZibo5JCl8/1iz0DuB6w902eHdn lMrUcXG4BVslojv8+MYc2hmtgxZgM4IuJpzOdGMRcok5blTJIiyCPr2H0G4nLrP6zvle7Gr9bcN Gq62a+uxlWhjk3CcL1pycwQr5YH39kDpv9OYWw8q24A==
X-Gm-Gg: ATEYQzwitURQzDSP5Rbmt5a7cugzY3OXTyjGTjoziAyAM1+/3j9D32nxiQBUj8hODRQ UjMjKbYwbrlygLMo+Q4qgn9h//YhrHk5dFnKXI5h6vKtwx38cesZ4oP81LeC9STRmOldAXFJ7N/ C4yodLCLAIjP/ff3al0M7tjMuGiAoCkXzWs4limImimlqn/ni5vlIvK2VEy3sjzNX7L3cKxaE5R NgUJiZO0Me+LuVBUBTbc8SBFkJxbmxajJj1ytd8Kwi9pF8dwu3zsm51Ueow+549tpwJbtiEf2Qz 9klcoN+5wA5Gm5pIRg==
X-Received: by 2002:a05:6402:4304:b0:65c:2377:3344 with SMTP id 4fb4d7f45d1cf-660f02d3aa6mr3264790a12.25.1772720199420; Thu, 05 Mar 2026 06:16:39 -0800 (PST)
MIME-Version: 1.0
References: <177248819212.3631768.8779340278624904983@dt-datatracker-6ff7c68975-7k42g> <CAOj+MMHbJXLoxDckdo-EF5-4FYv4g9PoopzM7z99rkX2kkbeAg@mail.gmail.com> <PH3PPFF8B8D687281529A6660726834C8FCC27FA@PH3PPFF8B8D6872.namprd11.prod.outlook.com> <CAOj+MMG5k86WZiDBe_tsBW_LvT4oLojU7wXruqGKgiXcxZ=wVQ@mail.gmail.com> <CAOj+MMHPEBhXbmPXMXHmCGAMQkP9jPPFVNAKS2sn=Yz1QzkMJA@mail.gmail.com> <PH3PPFF8B8D6872AB02E8DB7EAED4FF40ACC27FA@PH3PPFF8B8D6872.namprd11.prod.outlook.com> <CAOj+MMFTKsVw-V2S-EZZzJVvxp8Mdp26Hz4Pwrp9f5fjPSzg1w@mail.gmail.com> <174801dcabff$f8e43db0$eaacb910$@gmail.com> <CAEfhRrw3hC6si8td5SSW_MBcJ8KTNzfXjxwE0z=1z2d+wDBfEA@mail.gmail.com> <CAEfhRry5Y8ORg_pDePZhc-VN+CVNO7ATVMxiKoeRAyu+24aiLA@mail.gmail.com> <184a01dcac8a$7e80d930$7b828b90$@gmail.com> <CAOj+MMHmP1MbZ8SpPuv3QiRQOxidiCJDqXFWs3SKtM1FOfyuxQ@mail.gmail.com> <PH3PPFF8B8D6872A1DDC086E31208A30415C27DA@PH3PPFF8B8D6872.namprd11.prod.outlook.com> <CAOj+MMGL4s7eR3Muj9VYVRnmxDsbdVebJLP+f7p8R6UBCi8kTg@mail.gmail.com> <PH3PPFF8B8D6872DAE7AF7A71C9A786CAC1C27DA@PH3PPFF8B8D6872.namprd11.prod.outlook.com> <CAOj+MMFq7HDY1gK1jM5rKRO_zOnxK+ByJG5vX0h2ob2R0UFXGg@mail.gmail.com> <PH3PPFF8B8D6872D4984DEB272172761E94C27DA@PH3PPFF8B8D6872.namprd11.prod.outlook.com>
In-Reply-To: <PH3PPFF8B8D6872D4984DEB272172761E94C27DA@PH3PPFF8B8D6872.namprd11.prod.outlook.com>
From: Robert Raszuk <robert@raszuk.net>
Date: Thu, 05 Mar 2026 15:16:27 +0100
X-Gm-Features: AaiRm50Q9vDScY85RaRau4CDMV-bCnlx4eGS2MEb1CCZaxdEOPzF6c92G8A_sYg
Message-ID: <CAOj+MMFkNufFcFxCwTmkT85OFz25j875NxrH1NTCk6xMaDHELA@mail.gmail.com>
To: "Stephane Litkowski (slitkows)" <slitkows@cisco.com>
Content-Type: multipart/alternative; boundary="000000000000c5da89064c47953e"
Message-ID-Hash: FMFZQ4OB3NVTTX447QGPDGJY4F6LOB5D
X-Message-ID-Hash: FMFZQ4OB3NVTTX447QGPDGJY4F6LOB5D
X-MailFrom: robert@raszuk.net
X-Mailman-Rule-Misses: dmarc-mitigation; no-senders; approved; emergency; loop; banned-address; member-moderation; header-match-idr.ietf.org-0; nonmember-moderation; administrivia; implicit-dest; max-recipients; max-size; news-moderation; no-subject; digests; suspicious-header
CC: "Olivier Vroonen (ovroonen)" <ovroonen@cisco.com>, "kandhla.chandi@bell.ca" <kandhla.chandi@bell.ca>, "idr@ietf. org" <idr@ietf.org>
X-Mailman-Version: 3.3.9rc6
Precedence: list
Subject: [Idr] Re: I-D Action: draft-vroonen-idr-bgp-bestpath-nh-selection-01.txt
List-Id: Inter-Domain Routing <idr.ietf.org>
Archived-At: <https://mailarchive.ietf.org/arch/msg/idr/rTRU4Ts3r61wNWmVMPy9lToshJo>
List-Archive: <https://mailarchive.ietf.org/arch/browse/idr>
List-Help: <mailto:idr-request@ietf.org?subject=help>
List-Owner: <mailto:idr-owner@ietf.org>
List-Post: <mailto:idr@ietf.org>
List-Subscribe: <mailto:idr-join@ietf.org>
List-Unsubscribe: <mailto:idr-leave@ietf.org>

Hi,

> I think we are saying the same thing.

I am not sure ... You are clearly talking about situation where you ignore
next hop and instead use Tunnel Egress Endpoint address if say we consider
presence in the update message Tunnel Encapsulation Attribute as defined in
RFC9012.

Or worse you want both to get registered with corresponding RIB(s) and you
are asking for BGP to deal with both metrics one after the other one when
selecting the best path.

I don't think this will fly.

Also note that it does not work in most popular deployment model where it
is the next hop of service route recursively (see section 8 of RFC9012)
resolving via tunnel of your choice and update message itself with your
destination does not carry any tunnel info. And that is a much more
recommended deployment model from number of angles.

In fact RFC9012 clearly discourages use of egress endpoint different from
next hop:

During the deployment of techniques described in this document, operators
are encouraged to avoid mutually recursive route and/or tunnel
dependencies. There is greater potential for such scenarios to arise when
the tunnel egress endpoint for a given prefix differs from the address of
the next hop for that prefix.
Best,
R.








On Thu, Mar 5, 2026 at 1:06 PM Stephane Litkowski (slitkows) <
slitkows@cisco.com> wrote:

> Hi Robert,
>
> The key is the forwarding address, it may be the nexthop in most of the
> cases, but it could be something else like the tunnel endpoint of TEA or
> IPv6 address in SRV6 SID.
>
> BGP determines  using the information in the update what is the forwarding
> address to be used.
>
>
>
> The forwarding address (NH or something else) is registered to a table
> (any RIB or tunnel-table depending on implementation) and then the table
> notifies BGP about any change related to the registered forwarding address
> including the characteristics of the path (cost, source protocol, …)
>
>
>
> I think we are saying the same thing.
>
>
>
> > your configuration allowing or not given next hop to be
> considered valid for best path selection already proves that you can
> enforce NH validation base on the underlay preference of various forwarding
> paradigms.
>
>
>
> Saying pass/reject based on a criteria is not a preference as it will
> allow only some underlay paths to be used and some other will never be
> considered.
>
>
>
> Setting preference may look like:
>
>
>
> route-policy NH-SELECT
>
>   if protocol is isis 1 and route-type is level-2 and next-hop-type is
> mpls-te then
>
>    set nexthop-preference 200
>
>  elif  protocol is isis 1 and route-type is level-2
>
>  set nexthop-preference 100
>
> else
>
>     drop
>
>   endif
>
> end-policy
>
>
>
> In this case the pure IGP routes are still being considered by BGP, but
> they will be used only if there is no MPLS-TE path available at all
> otherwise we’ll primarly select a BGP path resolved over an MPLS TE tunnel.
>
>
>
> The way the preference is set is really implementation dependent, it could
> be from RPL as in the example, it could be inherited from RIB admin
> distance…
>
>
>
>
>
>
>
> Stephane
>
>
>
>
>
> *From:* Robert Raszuk <robert@raszuk.net>
> *Sent:* Thursday, March 5, 2026 12:43 PM
> *To:* Stephane Litkowski (slitkows) <slitkows@cisco.com>
> *Cc:* slitkows.ietf@gmail.com; Igor Malyushkin <gmalyushkin@gmail.com>;
> Olivier Vroonen (ovroonen) <ovroonen@cisco.com>; kandhla.chandi@bell.ca;
> idr@ietf. org <idr@ietf.org>
> *Subject:* Re: [Idr] Re: I-D Action:
> draft-vroonen-idr-bgp-bestpath-nh-selection-01.txt
>
>
>
> Hi,
>
>
>
> But your configuration allowing or not given next hop to be
> considered valid for best path selection already proves that you can
> enforce NH validation base on the underlay preference of various forwarding
> paradigms.
>
>
>
> But again this is mapped via next hop as it is key to BGP and it is NH
> which is used as a glue between BGP route and underlay transport.
>
>
>
> See you can select some path based on other best path selection step
> (existing or new) but it is still the next hop which will be passed to RIB
> (of yr choice inet.0 or inet.3 or inet.X etc ...) along with the service
> route and it is that NH which needs to be resolved to proper forwarding
> path.
>
>
>
> If you do not consider that I do not see a way how you glue BGP
> reachability with forwarding.
>
>
>
> Rgs,
>
> R.
>
>
>
>
>
> On Thu, Mar 5, 2026 at 12:34 PM Stephane Litkowski (slitkows) <
> slitkows@cisco.com> wrote:
>
> Hi Robert,
>
>
> Here is an example of XR config:
>
> route-policy NH-SELECT
>
>   if protocol is isis 1 and route-type is level-2 and next-hop-type is
> mpls-te then
>
>     pass
>
>   else
>
>     drop
>
>   endif
>
> end-policy
>
> !
>
> router bgp 1
>
> address-family ipv4 unicast
>
>   nexthop route-policy NH-SELECT
>
> !
>
> !
>
>
>
> Some from Juniper:
>
>
>
> policy-statement NH-SELECT {
>
>         term PERMIT {
>
>             from {
>
>                 protocol l-isis;
>
>                 rib inet.3;
>
>                 level 2;
>
>             }
>
>             then accept;
>
>         }
>
>         term REJECT {
>
>             then reject;
>
>         }
>
>     }
>
>
>
>
>
> I cannot disagree that you could play with metric offsets to mimic the
> preference but that is more a hack rather than a solution. Customer will
> have to compute what are the necessary offsets to add/substract depending
> on the minimum/maximum metric that we could observe for each type of
> underlay route, but there could still be a chance that it fails because the
> offset wasn’t good enough in some rerouting scenarios. Adding this
> “preference” is just more natural to use, less ambiguous and will always
> work. Again, it’s no more than what RIB does today when comparing two paths
> from different sources. Yes, it’s a change in the decision process, but it
> doesn’t change/impact existing deployments. People are free to use it or
> not.
>
>
>
>
>
> Stephane
>
>
>
>
>
> *From:* Robert Raszuk <robert@raszuk.net>
> *Sent:* Thursday, March 5, 2026 12:18 PM
> *To:* Stephane Litkowski (slitkows) <slitkows@cisco.com>
> *Cc:* slitkows.ietf@gmail.com; Igor Malyushkin <gmalyushkin@gmail.com>;
> Olivier Vroonen (ovroonen) <ovroonen@cisco.com>; kandhla.chandi@bell.ca;
> idr@ietf. org <idr@ietf.org>
> *Subject:* Re: [Idr] Re: I-D Action:
> draft-vroonen-idr-bgp-bestpath-nh-selection-01.txt
>
>
>
> Hi
>
>
>
> >  A lot of implementations already allow policies at NH resolution level.
>
>
>
> I am aware that some implementations do validate reachability to NH via
> RIB of your choice before declaring it is good or not. But as far as other
> policies I am not sure what would be some examples ? Can you provide some
> say cisco cli examples of such policies at NH resolution ?
>
>
>
> > Our point is that metric alone is not enough for BGP to select the most
> optimal path as you
>
> > could end up in situations where you compare two types of underlay
> routes that don’t
>
> > compute the cost in the same way (I agree that it’s not the mainstream
> and ideal scenario,
>
> > but it happens). So, BGP should have another tie breaker before the cost
> to make a fair
>
> > comparison happening. Where is “new” information is coming from:
>
>
>
> I fail to understand why metric is not enough. You may for example ask for
> adding 1000 to the metric of less preferred next hop. The end effect will
> be the same as introducing another best path selection step. Presumably
> before NH metric check.
>
>
>
> See selecting the next hop is only one part of it. You need to keep in
> mind that every time the underlay metric to next hop changes BGP gets a
> call back from RIB and needs to re-evaluate the paths (to check if perhaps
> there is no other better path available).
>
>
>
> Then also we do need to keep in mind consistent forwarding for networks
> which use encapsulation vs networks which do native hop by hop IP lookup.
>
>
>
> Thx,
>
> r
>
>
>
>
>
>
>
>
>
>
>
>
>
>
>
>
>
>
>
>
>
>
>
>
>
>
>
>
>
>
>
>
>
>
>
>
>
>
>
>
>
>
>
> On Thu, Mar 5, 2026 at 12:00 PM Stephane Litkowski (slitkows) <
> slitkows@cisco.com> wrote:
>
> Hi Robert,
>
> > But BGP has no clue about underlay path constraints
>
> Agree, but BGP may have information that allows to pick up the proper
> table in which to resolve the forwarding address: tunnel-type in TEA, color
> extcomm or whatever… I think we have an agreement on that.
>
>
> BGP may also have policies to prevent some specific types of transport to
> be used to resolve the NH (implicit or explicitly configured): for
> instance, in the BGP free core case that Igor mentioned, BGP could enforce
> that the forwarding address must be reachable using a tunneling technology.
> And if it’s not the case, it should consider NH as unreachable. A lot of
> implementations already allow policies at NH resolution level.
>
>
>
> Then from these tables will give to BGP the cost to reach the forwarding
> address.
>
>
>
> Our point is that metric alone is not enough for BGP to select the most
> optimal path as you could end up in situations where you compare two types
> of underlay routes that don’t compute the cost in the same way (I agree
> that it’s not the mainstream and ideal scenario, but it happens). So, BGP
> should have another tie breaker before the cost to make a fair comparison
> happening. Where is “new” information is coming from:
>
>    - It could also come directly from RIB (by using the admin distance)
>    - BGP NH policies could also be configured to set preferences based on
>    characteristics of the underlay route received from a RIB: protocol type,
>    is it tunnel or not, what type of tunnel… All this information should come
>    from RIB, as you mention, BGP cannot/must not guess them. It is just using
>    knowledge that it gets from RIB (but not limited just to the cost).
>
>
>
> Brgds,
>
>
>
> Stephane
>
>
>
>
>
> *From:* Robert Raszuk <robert@raszuk.net>
> *Sent:* Thursday, March 5, 2026 11:39 AM
> *To:* slitkows.ietf@gmail.com
> *Cc:* Igor Malyushkin <gmalyushkin@gmail.com>; Stephane Litkowski
> (slitkows) <slitkows@cisco.com>; Olivier Vroonen (ovroonen) <
> ovroonen@cisco.com>; kandhla.chandi@bell.ca; idr@ietf. org <idr@ietf.org>
> *Subject:* Re: [Idr] Re: I-D Action:
> draft-vroonen-idr-bgp-bestpath-nh-selection-01.txt
>
>
>
> Hi Stephan,
>
>
>
> You said:
>
>
>
> *If for instance, we cannot get a constrained path to N1 (and then we
> consider the IGP path to N1), but we have a constrained path to N2. Both
> underlay paths are not comparable anymore in term of costs. It’s first the
> “type” of underlay path that should tiebreak (prefer the constrained
> underlay path over non constrained). At this step, we need to introduce a
> small best path algo change and only in this particular scenario.*
>
>
>
> But BGP has no clue about underlay path constraints. How is this to be
> learned by BGP if you insist not to feed it to BGP with modified next hop
> metrics ?
>
>
>
> Note that what BGP carries is irrelevant as when it comes to any best path
> selection it will still only use the next hop when inserting the best path
> to RIB.
>
>
>
> So to reiterate ... if RIB (of your choice) can see path to both next hops
> properly it also can to return metric to such next hops properly to BGP for
> best path selection. BGP MUST not do any guessing here as it will put best
> path into such RIB and then RIB and FIB need to worry about it.
>
>
>
> You are making a mental shortcut - if we know that path to N2 is better we
> just select N2 ... but what we know at BGP is irrelevant - RIB must be of
> the same opinion for lot's of reasons.
>
>
>
> Many thx,
>
> Robert
>
>
>
>
>
> On Thu, Mar 5, 2026 at 11:26 AM <slitkows.ietf@gmail.com> wrote:
>
> Hi Igor,
>
> please find some comment inline.
>
> Brgds,
>
> Stephane
>
>
>
>
>
> *From:* Igor Malyushkin <gmalyushkin@gmail.com>
> *Sent:* Wednesday, March 4, 2026 8:40 PM
> *To:* slitkows.ietf@gmail.com
> *Cc:* Robert Raszuk <robert@raszuk.net>; Stephane Litkowski (slitkows) <
> slitkows@cisco.com>; Olivier Vroonen (ovroonen) <ovroonen@cisco.com>;
> kandhla.chandi@bell.ca; idr@ietf. org <idr@ietf.org>
> *Subject:* Re: [Idr] Re: I-D Action:
> draft-vroonen-idr-bgp-bestpath-nh-selection-01.txt
>
>
>
> Sorry. I lost the heading for L2VPN/L3VPN routes where the tunnel status
> is usually monitored and accounted for (the first bullet).
>
>
>
> ср, 4 мар. 2026 г. в 21:31, Igor Malyushkin <gmalyushkin@gmail.com>:
>
> Hi all,
>
> I have several thoughts on Section 2.1. I hope I got the idea right, at
> least, for some cases (I guess, there are several of them).
>
> Modern gear already works the way this draft wants it to. I.e., it does
> not install routes if there is no tunnel, and it monitors the tunnel's
> status, swapping the routes in Loc-RIB if there are ones to backup. I'd be
> really surprised to find anything behaving the other way. And nobody
> changed the best-path selection for that.
>
> [SLI] the draft addresses two things:
> - clarifies the additional or change in reachability checks that are
> required now and this is what ietf-idr-bgp-bestpath-selection-criteria
> started to do. As you mention, all/most of the implementations already do
> things right, the goal is just to align standard with the way
> implementations do. The draft will likely not create implementation
> changes. This part doesn’t require a best path selection change.
> - if nexthop/forwarding address is reachable and I come up with multiple
> iBGP paths that are equal, we need to compare the IGP cost to reach the NH.
> Two parts here:
>
> 1) As Robert mentioned, in an ideal case, you could expect that your
> forwarding endpoints are reachable using the same type of transport
> (including constraints used…), so we just need to pick up the right metric
> to consider (which may not be the IGP metric of the NH coming from default
> RIB). Again here, that doesn’t require any change in best path algo, only
> the source from which we pick the metric changes.
>
> 2) A small change in best path algo comes if for some reason you end up in
> having two iBGP paths for which forwarding addresses are not reachable
> through the same type of underlay path. Likely their costs are not
> comparable. You can still blindly compare, it will provide you a tie
> breaker but you could have chances to pick the wrong path. Such situation
> happens mainly when customers are asking for underlay path fallback. Let’s
> say you have Prefix P reachable through forwarding address N1 and N2 and
> customer wants to use a constrained path to reach N1 and N2, but if no
> constrained path exists then it is OK to follow a more traditional IGP
> based path. If for instance, we cannot get a constrained path to N1 (and
> then we consider the IGP path to N1), but we have a constrained path to N2.
> Both underlay paths are not comparable anymore in term of costs. It’s first
> the “type” of underlay path that should tiebreak (prefer the constrained
> underlay path over non constrained). At this step, we need to introduce a
> small best path algo change and only in this particular scenario.
>
>
> But let me stop on the other possible scenarios.
>
> *Unlabeled IPv4/v6 routes over MPLS transport.*
>
> For such routes, there may indeed be a problem when a transport is not
> ready yet, but the next-hop is reachable via the routing table. In the case
> of a BGP-free core, it is a dangerous situation, and I know of a couple of
> cases where it caused a lot of trouble. There are vendor knobs that prevent
> this behavior and cause routers to consider only the tunnel table, but
> people are not always aware of them or the issue in the first place. So,
> having something that tells us "I send you this precious route and, please,
> never use an unlabeled transport for it" is probably a good idea.
>
> [SLI] Absolutely agree, this caused a lot of issues in BGP free core
> network and it’s likely the reason why
> ietf-idr-bgp-bestpath-selection-criteria was initiated a long time ago.
> Most implementations have knobs or internal logic as you mention to handle
> that properly.
>
>
>
> *Labeled unicast IPv4/v6 routes.*
>
> This is a special case, and is too complicated. Vendors treat these routes
> differently. Some of them may consider these routes as vanilla ones,
> installing them into a routing table. Some of them allow filling the tunnel
> table if the next-hop is reachable over another tunnel. Some do both.
> Again, some vendors have magic knobs, although these knobs do not always
> allow you to achieve your goal.
>
>
>
> Considering the cases above, I do not see how the forwarding address (I
> mean exactly an address as an entity) fits, the forwarding address is
> always the same as the next-hop. Isn't it redundant? To me, it is enough to
> introduce some Sub-TLV in TEA instead.
>
> [SLI] I’m not following your point.
> Today in most of the scenarios, the forwarding address will still be the
> nexthop address. Cases where the forwarding address is different (and
> nexthop becomes useless) are:
> - tunnel-encaps (TEA has the endpoint address encoded so this is the one
> that we should use for reachability and cost retrieval and not the nexthop)
> - SRV6 services: we need to use the SID in the BGP prefix SID attribute
> and not the nexthop
> - SRTE like scenarios that use color or transport target that should use a
> combination of the nexthop address + something to get the forwarding
> address. Nexthop becomes for instance (color, NH) as lookup key: this case
> is really implementation dependent and nexthop could still be seen as the
> most important information for the lookup.
>
>
>
> My 2 cents.
>
>
>
> ср, 4 мар. 2026 г. в 19:55, <slitkows.ietf@gmail.com>:
>
> Hi Robert,
>
>
>
> Doing active/active, doesn’t always mean that any ingress should ECMP to
> the pool of active egress or should pick any path. Again, a lot of people
> rely on hot potatoe routing/shortest path (based on some criteria) to pick
> up the best egress. Picking “any” egress is not always acceptable.
>
>
>
>
>
> Brgds,
>
>
>
> Stephane
>
>
>
> *From:* Robert Raszuk <robert@raszuk.net>
> *Sent:* Tuesday, March 3, 2026 8:50 PM
> *To:* Stephane Litkowski (slitkows) <slitkows@cisco.com>
> *Cc:* Olivier Vroonen (ovroonen) <ovroonen@cisco.com>;
> kandhla.chandi@bell.ca; idr@ietf. org <idr@ietf.org>
> *Subject:* [Idr] Re: I-D Action:
> draft-vroonen-idr-bgp-bestpath-nh-selection-01.txt
>
>
>
> Hi,
>
>
>
> Ok glad we are on the same page on this. So far my impression was that the
> draft is trying to contain the solution within BGP and that drove most of
> the previous comments. So clearly it seems we both see a need (provided one
> is after this model) to go outside of BGP to get proper cost/metric.
>
>
>
> Now what is puzzling me even more is the lack of symmetry when advertising
> active-active service routes. See irrespective of the next hop used the
> transport should provide the optimal path to get there meeting application
> or service transport requirements. If so, selecting any exit for a service
> as best sould be ok. In fact while I very much do realize difficulties in
> doing so perhaps some considerations should be given to advertise such
> services with anycast next hops. Sync in application demux label/SID/uSID
> may be required though.
>
>
>
> Thx,
>
> R.
>
>
>
>
>
>
>
>
>
>
>
>
>
>
>
> On Tue, Mar 3, 2026 at 8:29 PM Stephane Litkowski (slitkows) <
> slitkows@cisco.com> wrote:
>
> Hi Robert,
>
> Right, the “underlay” transport is often not signaled/installed by BGP.
>
> What BGP needs is:
>
>    - what is the address to lookup (nexthop or something else)
>    - what is the “context” of this address , the context will tell where
>    BGP should get the resolution data from (is it from default RIB or from any
>    other instance, or in a tunnel table or anything…). For instance, if color
>    extcomm is used, the color may give a hint of where to resolve. If the BGP
>    path has a tunnel-encap telling that it’s a GRE tunnel, then again BGP
>    should know that it needs to track tunnels that ay reside somewhere else in
>    the system.
>    - Then BGP needs to do proper resolvability check in this context and
>    get proper “cost”
>
>
>
>
>
> *From:* Robert Raszuk <robert@raszuk.net>
> *Sent:* Tuesday, March 3, 2026 8:11 PM
> *To:* Stephane Litkowski (slitkows) <slitkows@cisco.com>
> *Cc:* Olivier Vroonen (ovroonen) <ovroonen@cisco.com>;
> kandhla.chandi@bell.ca; idr@ietf. org <idr@ietf.org>
> *Subject:* Re: I-D Action:
> draft-vroonen-idr-bgp-bestpath-nh-selection-01.txt
>
>
>
>
>
> Let me add something which I found perhaps not clearly communicated in my
> former messages.
>
>
>
> Your proposal calls for modification to best path selection to a
> metric/cost to a forwarding address instead of next hop. But various forms
> of transports may be installed outside of subject BGP AFI/SAFI in question.
> And my comment about going via RIB instance (inet.x) is based on the
> observation that recursion is happening there.
>
>
>
> If transport pipe is installed with controller .. or with distributed PCEP
> BGP SAFI which carries L3VPN or EVPN routes will not be aware of it. Such
> binding may and there often are done directly with proper next hop
> recursion or indirectly via some form of mapping (color). Again the
> transport information may very well reside and be signalled outside of BGP.
>
>
>
> Regards,
>
> Robert
>
>
>
> On Tue, Mar 3, 2026 at 2:55 PM Robert Raszuk <robert@raszuk.net> wrote:
>
> Hi Stephane,
>
>
>
> “Best” may not always mean from an IGP cost point of view.
>
>
>
> Understood. I am not talking about IGP cost ... but the "cost" or "metric"
> presented to BGP when BGP registers the interesting next hops to it. RIB
> can do all the magic it wants to normalize such cost and present to BGP as
> cost to next hop. It does not need to be an equal IGP SPF path at all.
>
>
>
> Two additional observations:
>
>
>
> 1)
>
>
>
> BGP allows you to use tie break in any point of insertion in the best path
> selection in the form of custom decision
> https://datatracker.ietf.org/doc/html/draft-ietf-idr-custom-decision-08
>
>
>
> 2)
>
>
>
> If you are offering a multihomed service and and are distributing paths
> with different next hops transport should be designed in a the correct way
> ... meaning required service path should be available to both next hops. If
> this is not the case then use existing tools like local pref to deprefer
> unwanted next hop.
>
>
>
> Best,
>
> R.
>
>
>
> On Tue, Mar 3, 2026 at 2:36 PM Stephane Litkowski (slitkows) <
> slitkows@cisco.com> wrote:
>
> Hi Robert,
>
>
>
> Thanks for your comment.
>
>
> Please see inline.
>
>
>
> Brgds,
>
>
>
> Stephane
>
>
>
>
>
> *From:* Robert Raszuk <robert@raszuk.net>
> *Sent:* Monday, March 2, 2026 11:22 PM
> *To:* Olivier Vroonen (ovroonen) <ovroonen@cisco.com>; Stephane Litkowski
> (slitkows) <slitkows@cisco.com>; kandhla.chandi@bell.ca
> *Cc:* idr@ietf. org <idr@ietf.org>
> *Subject:* Fwd: I-D Action:
> draft-vroonen-idr-bgp-bestpath-nh-selection-01.txt
>
>
>
> Hi,
>
>
>
> Few observations & a comment ...
>
>
>
> > A BGP update typically contains a prefix and a set of path attributes,
> including
>
> > the well-known mandatory NEXT_HOP attribute.
>
>
>
> Please kindly remove this. We are 19 years past publishing RFC4760 so we
> should strive to avoid further mess and confusion. NEXT_HOP Attribute is
> obsolete for all practical purposes and using such nomenclature is not
> helping anyone.
>
>
>
> > BGP NEXT_HOP attribute as defined in [RFC4271
> <https://www.ietf.org/archive/id/draft-vroonen-idr-bgp-bestpath-nh-selection-01.html#RFC4271>] and
> BGP NEXT_HOP field in the
>
> > MP_REACH_NLRI attribute as defined in [RFC4760
> <https://www.ietf.org/archive/id/draft-vroonen-idr-bgp-bestpath-nh-selection-01.html#RFC4760>] are
> referred to as NEXT_HOP
>
> > attribute in this document.
>
>
>
> Just say NEXT_HOP instead.
>
>
> [SLI] Agree, we’ll address it.
>
> - - - -
>
>
>
> Comment:
>
>
>
> It is nice that you included reference to
> Rajiv's draft-ietf-idr-bgp-bestpath-selection-criteria-12 draft. IMO it is
> pretty sufficient to make sure next hop is reachable.
>
>
>
> And IMO we should not go any further irrespective of choice of transport.
> What matters for BGP is if the path to next hop is valid or not. If you
> need to prefer one exit from the other use local preference within a
> domain. Do not overload NEXT_HOP tie break with your intradomain forwarding
> policies.
>
>
>
> Besides IGP metric to next hop is very far in best path selection check.
>
>
> [SLI] I agree that IGP metric to next hop is far in best path selection,
> however it is heavily used (Internet, LxVPN cases…) . Its position doesn’t
> reflect its importance.
>
> To your SR examples - it would be very inaccurate and in fact wrong to
> consider cost to first segment.
> [SLI] I agree, and this is actually not the case, in case of SR-TE path
> for instance, it’s the cost of the end-to-end path of SR policy (not the
> cost of the first segment). Is there something in the text that made you
> think that we should consider the cost of the first segment ?
>
>
>
> To summarize - I do not see that any changes are required in BGP. If you
> want to influence best path selection in BGP with your intradomain policies
> and various transport paradigmes expose that when you return to BGP a
> normalized metric to next hops in question.
>
> [SLI]  Few things:
>
>    - BGP has evolved since its early definition. We currently allow to
>    use an address which is no more the nexthop as the “next forwarder router”.
>    Thus, rules need to be adapted to keep the same logic as before, when I
>    have two iBGP paths that are “equal” from an attribute point of view, how
>    do I pickup the path for which the egress is best to me. “Best” may not
>    always mean from an IGP cost point of view.
>
>
>
>    - Metrics cannot really be normalized. The main issue comes when you
>    have two paths for which the “metric” is not comparable (e.g: an
>    accumulated delay vs an accumulated IGP cost based on other criteria).
>
>
>
>
>
> Your implementation has full freedom to do the right thing without
> touching BGP best path selection rules and without making it endlessly
> dependent on zoo of ever changing transport technologies.
>
> [SLI] The modification of the decision process is done:
>
>    - To ensure that we compare metric values that are comparable which is
>    IMO a good addition. Otherwise, we’ll tiebreak but we’ll likely not achieve
>    the goal of the IGP cost comparison step. We need to compare things that
>    are comparable. That’s the purpose of the “preference”, you can look at it
>    as the admin distance in RIB, RIB doesn’t compare costs of routes coming
>    from two protocols, they are not comparable.
>    - To correctly reflect that the nexthop may not be the source of the
>    “cost” value but another address may be used to reflect properly to which
>    next router we’ll forward to.
>
>
>
> Best,
>
> Robert
>
>
>
>
>
> ---------- Forwarded message ---------
> From: <internet-drafts@ietf.org>
> Date: Mon, Mar 2, 2026 at 10:50 PM
> Subject: I-D Action: draft-vroonen-idr-bgp-bestpath-nh-selection-01.txt
> To: <i-d-announce@ietf.org>
>
>
>
> Internet-Draft draft-vroonen-idr-bgp-bestpath-nh-selection-01.txt is now
> available.
>
>    Title:   BGP best path next-hop selection enhancements
>    Authors: Olivier Vroonen
>             Stephane Litkowski
>             Kandhla Chandi
>    Name:    draft-vroonen-idr-bgp-bestpath-nh-selection-01.txt
>    Pages:   18
>    Dates:   2026-03-02
>
> Abstract:
>
>    BGP [RFC4271] has originally been designed to carry IPv4 routing
>    information over the Internet.  IP routing being "hop-by-hop" in
>    nature, [RFC4271] defines the NEXT_HOP attribute which purpose is to
>    carry the address of the next router to send the IP packet to.  In
>    BGP, the next-hop may not be a directly connected router, hence, when
>    evaluating paths, a BGP speaker must determine if the next-hop is
>    resolvable and, if so, determine the internal cost to reach it.
>
>    The incremental use of tunneling technologies to carry traffic
>    between routers (e.g.: GRE, MPLS, SR-MPLS, SRv6...) may violate the
>    assumption that the address carried in the NEXT_HOP attribute is
>    representative of the actual forwarding next-hop.  These technologies
>    decouple the BGP control-plane's view of the next-hop from the data-
>    plane's actual forwarding endpoint.  This document describes the
>    problems that arise from this decoupling.  These problems include
>    sub-optimal path selection, incorrect resolvability tracking of the
>    forwarding path leading to traffic drop or misrouting, and others.
>    This document proposes some modification of BGP path selection
>    procedures to accommodate these use cases.
>
> The IETF datatracker status page for this Internet-Draft is:
>
> https://datatracker.ietf.org/doc/draft-vroonen-idr-bgp-bestpath-nh-selection/
>
> There is also an HTML version available at:
>
> https://www.ietf.org/archive/id/draft-vroonen-idr-bgp-bestpath-nh-selection-01.html
>
> A diff from the previous version is available at:
>
> https://author-tools.ietf.org/iddiff?url2=draft-vroonen-idr-bgp-bestpath-nh-selection-01
>
> Internet-Drafts are also available by rsync at:
> rsync.ietf.org::internet-drafts
>
>
> _______________________________________________
> I-D-Announce mailing list -- i-d-announce@ietf.org
> To unsubscribe send an email to i-d-announce-leave@ietf.org
>
> _______________________________________________
> Idr mailing list -- idr@ietf.org
> To unsubscribe send an email to idr-leave@ietf.org
>
>