[art] Re: Gdeflate as RFC

John R Levine <johnl@taugh.com> Tue, 17 September 2024 18:03 UTC

Return-Path: <johnl@taugh.com>
X-Original-To: art@ietfa.amsl.com
Delivered-To: art@ietfa.amsl.com
Received: from localhost (localhost [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 42423C14F6EE for <art@ietfa.amsl.com>; Tue, 17 Sep 2024 11:03:35 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.304
X-Spam-Level:
X-Spam-Status: No, score=-1.304 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, RCVD_IN_ZEN_BLOCKED_OPENDNS=0.001, RDNS_NONE=0.793, SPF_PASS=-0.001, T_SCC_BODY_TEXT_LINE=-0.01, T_SPF_HELO_TEMPERROR=0.01, URIBL_BLOCKED=0.001, URIBL_DBL_BLOCKED_OPENDNS=0.001, URIBL_ZEN_BLOCKED_OPENDNS=0.001] autolearn=no autolearn_force=no
Authentication-Results: ietfa.amsl.com (amavisd-new); dkim=pass (2048-bit key) header.d=iecc.com header.b="DiMedJcy"; dkim=pass (2048-bit key) header.d=taugh.com header.b="evv0Mv8M"
Received: from mail.ietf.org ([50.223.129.194]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id AC5cMDrIk33R for <art@ietfa.amsl.com>; Tue, 17 Sep 2024 11:03:30 -0700 (PDT)
Received: from gal.iecc.com (unknown [IPv6:2001:470:1f07:1126:0:43:6f73:7461]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature ECDSA (P-256) server-digest SHA256) (No client certificate requested) by ietfa.amsl.com (Postfix) with ESMTPS id D458EC14F5F5 for <art@ietf.org>; Tue, 17 Sep 2024 11:03:23 -0700 (PDT)
Received: (qmail 45479 invoked from network); 17 Sep 2024 18:03:21 -0000
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed; d=iecc.com; h=date:message-id:from:to:subject:in-reply-to:references:mime-version:content-type; s=b19966e9c469.k2409; t=1726596191; x=1726941791; bh=emd5I/sG1wy8LSF9iIOgWm8vh1EAuT5W706U2HGTRNI=; b=DiMedJcyqVmn81PvxdLTK4fwmH9zyPb4FnwSCrIbkMEiu/83ELyoO4K+sZBUqTEZkHS9FuRLpQEa+06jPnc4LnfY7lcR7pcw5z+O7tP4vAXvoG9bJj+Je3SI8dcWydZbAXeIg682QGoaXRFc0SCbDMvmy/vvYMVR/azwvIWUz9jpjIPR1UPzzrcwBajcsFUEx3MIh4fqsPEAZzrevLLfy7qYwIrt88M7nu67OAOwlXjzpxYqPDSlRbDKNpXjAzf8LFz6ArHjlLFm+CR6UFZHZoPFFRjNsywffHNYYdfRizIIElwdtj0gzVXrPokpsokQOJ8lEt+rgUd2XOi54Hv4aA==
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed; d=taugh.com; h=date:message-id:from:to:subject:in-reply-to:references:mime-version:content-type; s=b19966e9c469.k2409; bh=emd5I/sG1wy8LSF9iIOgWm8vh1EAuT5W706U2HGTRNI=; b=evv0Mv8MDGGf3zJBz8eXvm4F510Co2tl9RPgNw7Y9merd0yQApXO7omgaXizUu7B5WoIE4pG3yR/g2+bs86+sZK02RE9/KVDH0vKOi128NvQle2Tcrc5qQhQJGxabM5TZz1Rlz/wETtJ6lLVSzGtZLVpqPMUku8eVApQMXGRWFWjeVv8OZoH0X5jMbFmmVr5t5lI8eoy/XQ9zMrdwMAWTJodn8FF2+oMtLez+cUwosaruqiPW3wxsQ6Up4AboN8wq0Dqtqd1StBLIdhUlBxbgli19Ji2ox+MvlfRttf+z/v1aBnG/wRwAkOpgNrAHnzwYlHv9oStLDZe/8WqlLuTyw==
Received: from ary.qy ([IPv6:2001:470:1f07:1126::78:696d:6170]) by imap.iecc.com ([IPv6:2001:470:1f07:1126::78:696d:6170]) with ESMTPS (TLS1.3 ECDHE-RSA CHACHA20-POLY1305 AEAD) via TCP6; 17 Sep 2024 18:03:20 -0000
Received: by ary.qy (Postfix, from userid 501) id 7DB3F93CED60; Tue, 17 Sep 2024 14:03:19 -0400 (EDT)
Received: from localhost (localhost [127.0.0.1]) by ary.qy (Postfix) with ESMTP id DE80793CED42; Tue, 17 Sep 2024 14:03:19 -0400 (EDT)
Date: Tue, 17 Sep 2024 14:03:19 -0400
Message-ID: <8135f214-42c7-0346-aefa-433961ff9680@taugh.com>
From: John R Levine <johnl@taugh.com>
To: Vikram Kushwaha <vkushwaha@nvidia.com>, Akshay Subramaniam <asubramaniam@nvidia.com>, Martin Thomson <mt@lowentropy.net>, "art@ietf.org" <art@ietf.org>
X-X-Sender: johnl@ary.qy
In-Reply-To: <SJ1PR12MB6289D94D269F56930382302CC4612@SJ1PR12MB6289.namprd12.prod.outlook.com>
References: <SJ1PR12MB628929AD01C590E2D76D0736C4A82@SJ1PR12MB6289.namprd12.prod.outlook.com> <(vkushwaha=40nvidia.com@dmarc.ietf.org)<8734nq4vae.fsf@hobgoblin.ariadne.com> <SJ1PR12MB62896565C43EFFB259CC3EC0C4B12@SJ1PR12MB6289.namprd12.prod.outlook.com> <><20240801203828.BE49090BB693@ary.qy><8396945f-4fca-49b8-96af-817a48f9ebb6@betaapp.fastmail.com><8f3b7cdd-69ed-255a-174e-47abd1a05542@taugh.com><SJ1PR12MB6289C4CAF22C8FC155CFF010C4BF2@SJ1PR12MB6289.namprd12.prod.outlook.com>> <SJ1PR12MB62899575C7E5B464F8773ED5C4952@SJ1PR12MB6289.namprd12.prod.outlook.com> <f8bfd159-b153-4d88-b32f-283f7fa9f7ef@betaapp.fastmail.com> <SJ1PR12MB62899288451ED7DB39FB9567C4962@SJ1PR12MB6289.namprd12.prod.outlook.com> <SA3PR12MB90913F196825ABC042936E9BB9962@SA3PR12MB9091.namprd12.prod.outlook.com> <SJ1PR12MB6289D94D269F56930382302CC4612@SJ1PR12MB6289.namprd12.prod.outlook.com>
MIME-Version: 1.0
Content-Type: text/plain; charset="us-ascii"; format="flowed"
Message-ID-Hash: KV7XC2JBFXHYEQWFGWWQIYCACFYDVNMC
X-Message-ID-Hash: KV7XC2JBFXHYEQWFGWWQIYCACFYDVNMC
X-MailFrom: johnl@taugh.com
X-Mailman-Rule-Misses: dmarc-mitigation; no-senders; approved; emergency; loop; banned-address; member-moderation; header-match-art.ietf.org-0; nonmember-moderation; administrivia; implicit-dest; max-recipients; max-size; news-moderation; no-subject; digests; suspicious-header
X-Mailman-Version: 3.3.9rc4
Precedence: list
Subject: [art] Re: Gdeflate as RFC
List-Id: Applications and Real-Time Area Discussion <art.ietf.org>
Archived-At: <https://mailarchive.ietf.org/arch/msg/art/sU4rs4WClfI3tp-H4l4KBwHXi4A>
List-Archive: <https://mailarchive.ietf.org/arch/browse/art>
List-Help: <mailto:art-request@ietf.org?subject=help>
List-Owner: <mailto:art-owner@ietf.org>
List-Post: <mailto:art@ietf.org>
List-Subscribe: <mailto:art-join@ietf.org>
List-Unsubscribe: <mailto:art-leave@ietf.org>

Thanks for your note.  This is all informative but it still doesn't give 
me an understanding of why we would want to add this algorithm as an 
additional one for HTTP.

It's certainly faster in some circumstances but I'm not seeing how it'd 
offer an interesting difference on normal systems like PCs or phones that 
have a few CPU cores and a few GPU cores.

R's,
John

On Tue, 17 Sep 2024, Vikram Kushwaha wrote:

> Hi Martin/John,
> Please let us know if you have any more questions on GDeflate performance.
>
> Thanks!
>
> ~Vikram
>
> From: Akshay Subramaniam <asubramaniam@nvidia.com>
> Sent: Thursday, August 29, 2024 6:58 PM
> To: Vikram Kushwaha <vkushwaha@nvidia.com>; Martin Thomson <mt@lowentropy.net>; John R Levine <johnl@taugh.com>; art@ietf.org
> Subject: Re: [art] Re: Gdeflate as RFC
>
> Hi all,
>
> I can't read the full history of this email thread, especially plots or other attachments that were shared but let me add some thoughts on GDeflate vs ZSTD.
>
> I'm assuming the plots that you saw were some fairly coarse grained average compression ratio and throughput plots. Some clarifications on that:
>
>  1.  Both GDeflate and ZSTD in those plots are implemented on the GPU. Specifically, the ZSTD implementation used there is a custom CUDA based GPU implementation in the nvCOMP library, not the public CPU ZSTD library.
>
>  1.  The compression ratios for GDeflate are higher than ZSTD mainly because the compression techniques are different. Since GDeflate was designed for fast decompression, we implemented a high compression version of the compressor (similar to libdeflate level 12). Our GPU implementation of ZSTD was developed mainly for database applications where both compression and decompression throughput matter and so we tradeoff some compression ratio for throughput. The difference in compression ratio is only a result of the compressor implementations, not of the stream formats themselves.
>
>  1.  The speed difference between GDeflate and ZSTD depends a lot on the dataset. The main innovation in GDeflate is to make the entropy decoding much faster with the swizzling technique. But if the decompression performance is bottlenecked by LZ decompression, then the gap between GDeflate and ZSTD would be smaller.
>
>     *   Here's an example just for the Silesia corpus, GDeflate decompression throughput is 53 GB/s while ZSTD (on GPU) is 39 GB/s. So GDeflate is ~36% faster in this case. If we take a dataset that is much more entropy decode bound, the difference might be expected to grow.
> I want to reiterate that the ZSTD numbers are from a highly optimized GPU implementation in nvCOMP. If you look at the performance of the regular ZSTD libary on lzbench<https://github.com/inikep/lzbench>, it tops out at 1.2GB/s.
>
> There is one other use case for GDeflate that might be interesting. Since the main innovation in GDeflate is from faster entropy coding, we can use a very fast throughput compression algorithm by only entropy coding data without the LZ phase. This allows for symmetric compression and decompression throughputs of 150-200 GB/s and allows for compression to be used in communication bound applications where communication is done over relatively high performance networks.
>
> Thanks,
> Akshay
> ________________________________
> From: Vikram Kushwaha <vkushwaha@nvidia.com<mailto:vkushwaha@nvidia.com>>
> Sent: Thursday, August 29, 2024 10:36 AM
> To: Martin Thomson <mt@lowentropy.net<mailto:mt@lowentropy.net>>; John R Levine <johnl@taugh.com<mailto:johnl@taugh.com>>; art@ietf.org<mailto:art@ietf.org> <art@ietf.org<mailto:art@ietf.org>>; Akshay Subramaniam <asubramaniam@nvidia.com<mailto:asubramaniam@nvidia.com>>
> Subject: RE: [art] Re: Gdeflate as RFC
>
> I am not familiar with ZSTD, but adding @Akshay Subramaniam who will be able to better answer questions on the ZSTD vs Deflate performance.
>
>
> ~Vikram
>
> -----Original Message-----
> From: Martin Thomson <mt@lowentropy.net<mailto:mt@lowentropy.net>>
> Sent: Wednesday, August 28, 2024 8:52 PM
> To: Vikram Kushwaha <vkushwaha@nvidia.com<mailto:vkushwaha@nvidia.com>>; John R Levine <johnl@taugh.com<mailto:johnl@taugh.com>>; art@ietf.org<mailto:art@ietf.org>
> Subject: Re: [art] Re: Gdeflate as RFC
>
> Thanks for sharing that Vikram, it's helpful.
>
> I'm curious as to what zstd folks think about these results.  Is it the case that the tuning of zstd was to ensure that the throughput would be comparable to gdeflate?  From my understanding, zstd is capable of far better compression ratios than deflate and that your work was primarily focused on throughput, such that gdeflate would necessarily outperform zstd on that axis as much as it seems to have done in this scenario.
>
> I'm not sufficiently expert here, so I'll defer to others.  However, those numbers aren't necessarily convincing (though the OOMs on zstd might be if it were explained).  zstd seems to have slightly less throughput and compression both, but it's not so clearly a win as I'd have expected.  LZ4 is clearly unsuitable, but that's expected.
>
> On Thu, Aug 29, 2024, at 06:29, Vikram Kushwaha wrote:
>> Hi John/Martin,
>>
>> As I continue to work on the IPR disclosures, as requested, attaching
>> comparisons of GDeflate performance with other compression algorithms.
>> The summary is that GDeflate has a high throughput while maintaining a
>> high compression ratio.
>>
>> Thanks,
>> ~Vikram
>>
>> -----Original Message-----
>> From: Vikram Kushwaha <vkushwaha@nvidia.com<mailto:vkushwaha@nvidia.com>>
>> Sent: Monday, August 5, 2024 9:51 PM
>> To: John R Levine <johnl@taugh.com<mailto:johnl@taugh.com>>; Martin Thomson
>> <mt@lowentropy.net<mailto:mt@lowentropy.net>>; art@ietf.org<mailto:art@ietf.org>
>> Subject: RE: [art] Re: Gdeflate as RFC
>>
>> Thanks all for your feedback.
>>
>> I am still reading up on the IP disclosures, I will follow up on that
>> in an another email.
>>
>>> What sort of applications do you have in mind here?
>> This will be useful in applications where GPU decompression can be
>> done on the fly, offloading the work from CPU. One use case is gaming
>> where frames are streamed and data asset decompressions are handled by
>> some of the GPU cores.
>>
>>
>>> If it's supposed to be generally useful it'd also be helpful to have some idea how it works on normal CPUs.
>> That would depend on how the decompressor is written. If a CPU based
>> decompressor can make use of the parallel nature of the decompression
>> with gdeflate(it should be with thread programming) it will be faster,
>> though we haven't written one and so do not have numbers for
>> comparison as CPU cores vary a lot.
>>
>> I am going to send out an email comparing gdeflate with zstd and other formats.
>>
>> Of course, lower end GPUs will show a smaller gain but it will still
>> be faster than CPU decompression. The advantage we are going for is by
>> offloading decompression to GPU, CPU can be used for other tasks. This
>> is quite useful in gaming/visualization apps where CPU to GPU
>> communication is a bottleneck.
>>
>> One more thing I would like to add is that Microsoft has already
>> adopted GDeflate as their default GPU decompression method, so by
>> promoting this as a RFC we were hoping to standardize it more globally.
>>
>> ~Vikram
>>
>> -----Original Message-----
>> From: John R Levine <johnl@taugh.com<mailto:johnl@taugh.com>>
>> Sent: Thursday, August 1, 2024 8:29 PM
>> To: Martin Thomson <mt@lowentropy.net<mailto:mt@lowentropy.net>>; art@ietf.org<mailto:art@ietf.org>
>> Cc: Vikram Kushwaha <vkushwaha@nvidia.com<mailto:vkushwaha@nvidia.com>>
>> Subject: Re: [art] Re: Gdeflate as RFC
>>
>> On Fri, 2 Aug 2024, Martin Thomson wrote:
>>> On Fri, Aug 2, 2024, at 06:38, John Levine wrote:
>>>> What sort of applications do you have in mind here? The main place
>>>> DEFLAATE is used in IETF protocols is in compressed HTTP streams
>>>> which doesn't strike me as the kind of thing one would often do on a GPU.
>>>
>>> I don't see why not, if the GPU is that much faster.
>>
>> That benchmark is on an RTX3090 which has 10,000 cores.  I'm typing
>> this on an M2 Pro laptop whose GPU has 19 cores.  I believe their
>> benchmark but I also don't see how it's relevant to anything but high
>> end dedicated GPUs.
>>
>>> I'm more concerned about the comparison to more modern stuff, like
>>> brotli or zstd.  Those tend to offer far better compression
>>> performance than deflate. ...
>>
>> That is an excellent point if we're going to change the comprsssion scheme.
>>
>> R's,
>> John
>>
>>
>> Attachments:
>> * GDeflateComparison.pdf
>

Regards,
John Levine, johnl@taugh.com, Taughannock Networks, Trumansburg NY
Please consider the environment before reading this e-mail. https://jl.ly