Re: [New Issue] CID based attacks

Kashyap Thimmaraju <k.thimmaraju@informatik.hu-berlin.de> Wed, 25 November 2020 13:30 UTC

Return-Path: <k.thimmaraju@informatik.hu-berlin.de>
X-Original-To: quic@ietfa.amsl.com
Delivered-To: quic@ietfa.amsl.com
Received: from localhost (localhost [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 45E693A12AF for <quic@ietfa.amsl.com>; Wed, 25 Nov 2020 05:30:02 -0800 (PST)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -2
X-Spam-Level:
X-Spam-Status: No, score=-2 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, NICE_REPLY_A=-0.001, SPF_HELO_NONE=0.001, SPF_PASS=-0.001, URIBL_BLOCKED=0.001] autolearn=ham autolearn_force=no
Authentication-Results: ietfa.amsl.com (amavisd-new); dkim=pass (1024-bit key) header.d=informatik.hu-berlin.de
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id 6681bw2Trr3l for <quic@ietfa.amsl.com>; Wed, 25 Nov 2020 05:30:00 -0800 (PST)
Received: from mailout1.informatik.hu-berlin.de (mailout1.informatik.hu-berlin.de [141.20.20.101]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by ietfa.amsl.com (Postfix) with ESMTPS id BB8883A12B3 for <quic@ietf.org>; Wed, 25 Nov 2020 05:29:58 -0800 (PST)
Received: from mailbox.informatik.hu-berlin.de (mailbox [141.20.20.63]) by mail.informatik.hu-berlin.de (8.15.1/8.15.1/INF-2.0-MA-SOLARIS-2.10-25) with ESMTPS id 0APDTuqa013418 (version=TLSv1.2 cipher=DHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK) for <quic@ietf.org>; Wed, 25 Nov 2020 14:29:56 +0100 (MET)
Received: from hashkash-imac.local ([87.123.148.220]) (authenticated bits=0) by mailbox.informatik.hu-berlin.de (8.15.1/8.15.1/INF-2.0-MA-SOLARIS-2.10-AUTH-26-465-587) with ESMTPSA id 0APDTtSV013415 (version=TLSv1.2 cipher=AES256-GCM-SHA384 bits=256 verify=NO) for <quic@ietf.org>; Wed, 25 Nov 2020 14:29:56 +0100 (MET)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=informatik.hu-berlin.de; s=mailbox; t=1606310996; bh=/T/nw20IwTsaeolZEqzWDry7bcBiL9oJ8GQlyvthSxA=; h=Subject:From:To:References:Date:In-Reply-To; b=qPe/6Yj+rqMkj54OHO89LPWKoPhzipfk4IbGo8dWbxONENaFCh699Lq8xRK6nmvFk Q9IM80Ipp7pFhnMj7OIxg8ix8bTLwz4X1ImNS8GKQfSVgDbNoAbXC4EJIoE2KK7cs1 xJ68vMbnJIburbIVXoH6RssREl2+bfhMDW9u5s/Q=
Subject: Re: [New Issue] CID based attacks
From: Kashyap Thimmaraju <k.thimmaraju@informatik.hu-berlin.de>
To: quic@ietf.org
References: <7f6e6106-60d6-8e9d-4566-b5e115099e9b@informatik.hu-berlin.de>
Message-ID: <e0e72b88-b9a0-e772-348a-92f23101643b@informatik.hu-berlin.de>
Date: Wed, 25 Nov 2020 14:29:55 +0100
User-Agent: Mozilla/5.0 (Macintosh; Intel Mac OS X 10.15; rv:78.0) Gecko/20100101 Thunderbird/78.5.0
MIME-Version: 1.0
In-Reply-To: <7f6e6106-60d6-8e9d-4566-b5e115099e9b@informatik.hu-berlin.de>
Content-Type: text/plain; charset="utf-8"; format="flowed"
Content-Transfer-Encoding: 8bit
Content-Language: en-US
X-Virus-Scanned: clamav-milter 0.102.3 at mailbox
X-Virus-Status: Clean
X-Greylist: Sender succeeded STARTTLS authentication, not delayed by milter-greylist-4.6.1 (mail.informatik.hu-berlin.de [141.20.20.50]); Wed, 25 Nov 2020 14:29:56 +0100 (MET)
Archived-At: <https://mailarchive.ietf.org/arch/msg/quic/1cGL5QazZGjybhyIWjYF0O684iA>
X-BeenThere: quic@ietf.org
X-Mailman-Version: 2.1.29
Precedence: list
List-Id: Main mailing list of the IETF QUIC working group <quic.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/quic>, <mailto:quic-request@ietf.org?subject=unsubscribe>
List-Archive: <https://mailarchive.ietf.org/arch/browse/quic/>
List-Post: <mailto:quic@ietf.org>
List-Help: <mailto:quic-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/quic>, <mailto:quic-request@ietf.org?subject=subscribe>
X-List-Received-Date: Wed, 25 Nov 2020 13:30:02 -0000

Hi QUIC-WG,

Thanks for recognizing my work and effort.

I was not part of the mailing list (but I am now :)) when you replied
so I've had to manually copy your responses. Below I've included the
person's name and replied to their comments/remarks.

 > Dmitri Tikhonov
 > I ran some experiments myself and I realized that my original guess was
 > incorrect.  lsquic server retires the Initial DCID a short time after
 > handshake succeeds.  When a CID is retired, all incoming packets bearing
 > it are dropped.
 >
 > lsquic keeps CIDs retired for at least 30 seconds and at most forever
 > (old entries are purged opportunistically when new entries are added.)

Good to know. We did some experiments and it seemed to persist for over 
a day.
However, we did not try to connect with other CIDs so that could explain why
they remained for so long.

 > I looked for guidance in the Transport Draft for how long a CID is to
 > stay retired (both Initial DCID and those retired via the
 > RETIRE_CONNECTION_ID frame), but found none.

So for lsquic, the problem is about a CID being stuck in the retired state?

 > Christian Huitema
 > The description of the experiment does not say whether the successive
 > connection attempts used the same IP address. For Picoquic at least
 > that's important, because Picoquic retrieves the handshake context using
 > the combination of Initial DCID and client IP + port. Multiple
 > connection attempts using the same Initial DCID and different IP
 > addresses will be treated as independent connections. This was one of
 > the suggestion in
 > https://tools.ietf.org/html/draft-kazuho-quic-authenticated-handshake-01.

The same IP address was used for the successive connection. However, the 
source
port was different.

 > Kazuho Oku
 > Quicly adopts the same approach. Applications of quicly are expected to
 > supply their own CID generation scheme, and therefore quicly does not 
know
 > if there's enough entropy in the CID to avoid collision between
 > server-supplied CIDs and the original DCID being generated by the client.
 >
 > Therefore, during the handshake, quicly uses `4-tuple && 
(client-generated
 > DCID || server-generated CID)` as the packet routing scheme.
 >
 > That's how we avoid the problem raised by Kashyap.

If the client-generated DCID is used, why did I not see the problem with 
quicly?
Ah, because the source port in the 2nd attempt is different from the first.

 > Martin Thomson
 > Firstly, thanks for taking a look at this.  This is obviously a 
considerable
 > amount of work and it is good to see people thinking about the way 
that the
 > different pieces fit together.

You're welcome.

 > That isn't the whole story though.  If the *routing* infrastructure 
does the
 > same thing, then you are able mount the claimed attack by varying source
 > address.  If the routing infrastructure only looks at the connection 
ID, then
 > you won't reveal any information.  But either allows targeting of server
 > instances, which is probably unwise, so I would be surprised to see 
that in
 > more advanced infrastructure.  Address tuple-based routing, which 
might still
 > be common, does offer some opportunities here, but that was already 
true as
 > those systems can be exploited by manipulating source address.

The finer granularity in the routing infrastructure does influence the
feasibility of the attack at a first glance. However, if CID routing is 
used, I
don't think the attack is mitigated. AFAIU, CID routing will be based on the
server generated CID. Hence, if the behavior described exists, and 
assuming the
load balancer uses some kind of hashing to spray requests across the 
instances,
the attacker could eventually enumerate the instances.

 > All in all, the question of how a load balancer directs new 
connections to
 > server instances is highly relevant here.
Exactly.

 > It probably pays to see what an attacker gains.  Revealing what 
connection IDs
 > are in use is something, but given the size of the space, that might 
not be
 > especially valuable.  And exploiting this requires a routing 
infrastructure
 > that is vulnerable to more interesting attacks, like resource 
exhaustion by
 > targeting.

I think the point is not whether the CID is revealed. It is that an attacker
can build on this to eventually count the number of reachable instances. 
I did
not include that part of my research here as we are still working on it.
However, I do have a sketch of how it could be done below.

Enumeration Algorithm. We repeatedly issue connection requests with a
sequentially increasing source port, the CIDs however remain the same 
for each
request. We then count the number C of successful connections 
established. If we
do not receive a response from the server, we have reached an instance 
that was
already counted. If RR is being used, then it could be that other clients’
requests interleaved ours, hence, resulting in our request reaching a 
previously
seen instance again instead of an uncounted one. If hashing is being 
used, then
it could be that our 4-tuple values (source IP and port, and destination 
IP and
port) hashed to the same value, and hence the same instance. Therefore, we
continue to issue further requests until we do not establish any new 
connection
with the server after a threshold of max requests attempts. After which 
C will
be the number of server instances.

 > The covert channel is something we've already decided is not 
interesting.  The
 > number of other covert channels is practically unbounded here, and 
the bit rate
 > of the one you describe is far lower.

Okay. I agree the throughput is not the highest, and I have not tested 
it over
the WAN. Nevertheless, I see it being used in a positive way, e.g., to
circumvent censorship. However, it could also be used as some kind of 
rendezvous
protocol, e.g., for bots to signal to a C&C server or prepare to mount an
attack. The root cause is the behavior I described in my first email.

 > Jana Iyengar
 > Thank you, Kashyap, for doing this work and for bringing it to the
 > working group.

You're welcome.

 > I agree with Martin's assessment, specifically that the interesting 
exploit
 > on enumerating the number of servers, is very dependent on how load
 > balancing is done. A relevant point here is that the attacker can only
 > control the DCID in the first flight of Initial packets, and if the 
server
 > treats Initial packets differently than it does subsequent packets, 
that is
 > yet another way in which the surface of this exploit gets limited.

Thank you for appreciating my work and discussing the matter so openly.  
I am
however not fully convinced by the argument made by Martin. True, there are
mentions of operational and management guidance as well as the quic-lb 
draft.
However, I believe, the key observation I made has been overlooked, 
which is the
unspecified behavior in the spec. when it comes to dealing with the same 
DCID
across successive connections. As my tests have shown, there is a 
difference in
the way implementations handle such a scenario which makes the enumeration
attack feasible on at least 4 different implementations. Also, the point 
raised
by Dimitri may also need some attention, i.e., wrt the CID retire timeout.

Sincerely,

-- 
Dr.-Ing Kashyap Thimmaraju

Lehrstuhl für Technische Informatik
Institut für Informatik
Humboldt-Universität zu Berlin

Besucheranschrift:
Rudower Chaussee 25, 12489 Berlin
Haus 4, 3. OG

k.thimmaraju@informatik.hu-berlin.de

http://www.ti.informatik.hu-berlin.de