[art] Re: Gdeflate as RFC

Vikram Kushwaha <vkushwaha@nvidia.com> Tue, 17 September 2024 18:09 UTC

Return-Path: <vkushwaha@nvidia.com>
X-Original-To: art@ietfa.amsl.com
Delivered-To: art@ietfa.amsl.com
Received: from localhost (localhost [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 97500C14F6BC for <art@ietfa.amsl.com>; Tue, 17 Sep 2024 11:09:49 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -2.254
X-Spam-Level:
X-Spam-Status: No, score=-2.254 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIMWL_WL_HIGH=-0.148, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, RCVD_IN_MSPIKE_H2=-0.001, RCVD_IN_ZEN_BLOCKED_OPENDNS=0.001, SPF_NONE=0.001, T_SCC_BODY_TEXT_LINE=-0.01, URIBL_BLOCKED=0.001, URIBL_DBL_BLOCKED_OPENDNS=0.001, URIBL_ZEN_BLOCKED_OPENDNS=0.001] autolearn=ham autolearn_force=no
Authentication-Results: ietfa.amsl.com (amavisd-new); dkim=pass (2048-bit key) header.d=nvidia.com
Received: from mail.ietf.org ([50.223.129.194]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id qSA4T79ZEqny for <art@ietfa.amsl.com>; Tue, 17 Sep 2024 11:09:45 -0700 (PDT)
Received: from NAM11-BN8-obe.outbound.protection.outlook.com (mail-bn8nam11on2086.outbound.protection.outlook.com [40.107.236.86]) (using TLSv1.2 with cipher ECDHE-ECDSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by ietfa.amsl.com (Postfix) with ESMTPS id 448BDC14F69C for <art@ietf.org>; Tue, 17 Sep 2024 11:09:45 -0700 (PDT)
ARC-Seal: i=1; a=rsa-sha256; s=arcselector10001; d=microsoft.com; cv=none; b=jxXdJAhgjH3Hygunn3EHAASELUtrflQ1weNDNKs2UrBdvZ+Y8DLLCjhioqrqhBNZfDYKR3DsTwfz/ihKlzVABJtJs5/R4afxpVOD8R51S1h2o0o/uIgEBupgSHBqFzFzGVtbWQHp9xWTEVI2turFtsbxuEUJBhax9YXtim0vg8HbiDfP+uBf1chpJKv7+ZPMPjIzhALKLvdTrBUuUff8k0wXI6hXCRW9ZD4vF9zv+lE3mubnTzIAzOyU3400lMPAGAEATHGqY6Rls0Lu4c3esgh0mUc864QxSUe8Ud2T/FZ0+43oqliwqNsMVBqAJ88t5ipBPDK9ZIWu77IeB0EPJA==
ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector10001; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-AntiSpam-MessageData-ChunkCount:X-MS-Exchange-AntiSpam-MessageData-0:X-MS-Exchange-AntiSpam-MessageData-1; bh=swA5G4ANiQnXpFbVfW2Du1Hv0a7WCLMs1wJI7Lsr3nU=; b=wiimMu8aSrUukJPu1Kq4QA+zNcqG40xEkEog9i5UyGrKCeYyF7o9HaKddQozSxashsVzfGfyB3Au1qsLu9Vog3VIpfzJT8hJ2bm/S3d4B98WFZZSZxiE34g0nXVoN1Avn1HeRp6FjwfMv/drU39jqcgIU4SZ1lGTswH9cbbX0wx0OyhQqDym8sjRUHGHhYoJiPLpomGdl92pRUXrBWa3bKCyUuPuQRzNAMUx7AS0fZ/K5oqjJ5Wm+3lqdqTvP1zDdvPFzPBJ41EaG8ob78PvmrHBQaobmckJs0CcqUj+b3HGYzjnsDA8rgBLRsZ5ecJ+zPGYnJaqXBlc+7j9dJWOhA==
ARC-Authentication-Results: i=1; mx.microsoft.com 1; spf=pass smtp.mailfrom=nvidia.com; dmarc=pass action=none header.from=nvidia.com; dkim=pass header.d=nvidia.com; arc=none
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=Nvidia.com; s=selector2; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-SenderADCheck; bh=swA5G4ANiQnXpFbVfW2Du1Hv0a7WCLMs1wJI7Lsr3nU=; b=PG8WDi7S+sqJ1bJMYzn5TJQCMfTVx7a6sK6fe/9AUk1w9Y+AfD5RshVfgsIAFODyI+2Yf67R2A44d9aKvytab0M8KboFazzQq3kF0I1GZKeTiQvH1gJSu1+7Ga2nkkjC/XLL+XkNeX5ADjqqHSm0/+w23BsylPSjx8ZacVwkN+ygMkayESncZzuCjD4dp8UYFRM3Q1ka8CONpviCsMncLLlxaNU3+he/cc2ZdybKG8X3K9Jc5Ul5R7Wc9yPh/PNjOukfAacqDaYR2BgtuFxBjx984yHkkSV9XE3OGfOUiXq5o5cAgt1mmLQ1BndKvUapgDcUvbWnFIy9a7P52j529A==
Received: from SJ1PR12MB6289.namprd12.prod.outlook.com (2603:10b6:a03:458::17) by DS0PR12MB7969.namprd12.prod.outlook.com (2603:10b6:8:146::19) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.20.7962.24; Tue, 17 Sep 2024 18:09:40 +0000
Received: from SJ1PR12MB6289.namprd12.prod.outlook.com ([fe80::57a7:c49a:cfb1:3be3]) by SJ1PR12MB6289.namprd12.prod.outlook.com ([fe80::57a7:c49a:cfb1:3be3%4]) with mapi id 15.20.7962.022; Tue, 17 Sep 2024 18:09:40 +0000
From: Vikram Kushwaha <vkushwaha@nvidia.com>
To: John R Levine <johnl@taugh.com>, Akshay Subramaniam <asubramaniam@nvidia.com>, Martin Thomson <mt@lowentropy.net>, "art@ietf.org" <art@ietf.org>
Thread-Topic: [art] Re: Gdeflate as RFC
Thread-Index: AQHa5FLb4IuvPEdxVEGeQ4uSUkGTDrIS+RkAgAAkJwCABl90oIAjyeUQgABMDgCAARf7UIAAVjhagB1Y0lCAADVvgIAAAFNg
Message-ID: <SJ1PR12MB62897D144A600FF377EE095DC4612@SJ1PR12MB6289.namprd12.prod.outlook.com>
References: <SJ1PR12MB628929AD01C590E2D76D0736C4A82@SJ1PR12MB6289.namprd12.prod.outlook.com> <(vkushwaha=40nvidia.com@dmarc.ietf.org)<8734nq4vae.fsf@hobgoblin.ariadne.com> <SJ1PR12MB62896565C43EFFB259CC3EC0C4B12@SJ1PR12MB6289.namprd12.prod.outlook.com> <><20240801203828.BE49090BB693@ary.qy><8396945f-4fca-49b8-96af-817a48f9ebb6@betaapp.fastmail.com><8f3b7cdd-69ed-255a-174e-47abd1a05542@taugh.com><SJ1PR12MB6289C4CAF22C8FC155CFF010C4BF2@SJ1PR12MB6289.namprd12.prod.outlook.com>> <SJ1PR12MB62899575C7E5B464F8773ED5C4952@SJ1PR12MB6289.namprd12.prod.outlook.com> <f8bfd159-b153-4d88-b32f-283f7fa9f7ef@betaapp.fastmail.com> <SJ1PR12MB62899288451ED7DB39FB9567C4962@SJ1PR12MB6289.namprd12.prod.outlook.com> <SA3PR12MB90913F196825ABC042936E9BB9962@SA3PR12MB9091.namprd12.prod.outlook.com> <SJ1PR12MB6289D94D269F56930382302CC4612@SJ1PR12MB6289.namprd12.prod.outlook.com> <8135f214-42c7-0346-aefa-433961ff9680@taugh.com>
In-Reply-To: <8135f214-42c7-0346-aefa-433961ff9680@taugh.com>
Accept-Language: en-US
Content-Language: en-US
X-MS-Has-Attach:
X-MS-TNEF-Correlator:
authentication-results: dkim=none (message not signed) header.d=none;dmarc=none action=none header.from=nvidia.com;
x-ms-publictraffictype: Email
x-ms-traffictypediagnostic: SJ1PR12MB6289:EE_|DS0PR12MB7969:EE_
x-ms-office365-filtering-correlation-id: 7ca2727c-dc4c-4338-e7e6-08dcd743e654
x-ms-exchange-senderadcheck: 1
x-ms-exchange-antispam-relay: 0
x-microsoft-antispam: BCL:0;ARA:13230040|376014|366016|1800799024|38070700018;
x-microsoft-antispam-message-info: 5OSAN5avzgEsbmhXbrJdF2zoFk+WJpdTAAEM33i0hdDB46fi7ee71r6k45b2m7sHyWZx++ZpG5xt6h2sIHuMz/2E/tOxfddqqoqZCN3MGTmg+Oj2shONtJQM751qirXiUznpB3JqiudBc15KzL6j+ZQJVJNFxGlAnCeatbnHBOQvWZV/XR3WqWsWbCaMVn6oBvNYj6Y5wrHCvm7538JeES34uqhnCX4FfZuopfYZpSp2MQXRVPbjSFJWiOGB34be1Ew0sqHnr91p98tHAsH8oJkasMW6dldoRTczY7lgJ4lzO30tj3SDvh+W7SowrSSv6NP8R5hRPOVAH2yQ7E9iMQXXp0YPvaKK1lGBoPrD5ZLAV8hEBeK3yLR+6r6y8p+1HTIUy15wZh6h6QgubORgfTn1vo5Vya8h81lqUawN3lHBCYdLHCstNu/QABxXn2OKblk/74/UxA8gp5312ch5hNxrtDJ2ISzkfO7ZSnLsenULwzH39nEvCNWRYB+q/+c3IJpWojle7mLrY0a6uw82bdKiF4J4MibXXtz5YbH1IYg4er+Es5kAoYZrj8V2rxzGTPbOQu1NPK7ebac1ouT/PWUdpAHvPuGFJmduH45okvGEnef7tF6NBQ1FF97/8iansMZKwFH9M5NGCa8KbbjUXJ5PBs0aTAxhor6XjPcJxBq+1ONJjTere35J4dCyb/R5MMnhTbkvo0mY5d3JD1NI5jLWGPvmdwc5n9rWQCH9rFktlW/KwoGLycgseV7wjG92/LF4UZjP31ZxtaFQNk/1OluMQV0+OLEhpdANfvNvcHGJO7rwHTdYRMbLIwY2noywvHgP0YSneI2IfDt46wfmLx0PbtDbjZg1iLKnwqVzKAmVDTh+rFSPTH70hu/IgT6uXSi4p4l+hKe5L2QElHq/8CbIsTcqdJpMr0V/VY4Wecl0D/gifJbASJZY1WDkIP74gmZQSkI8L4PRL3S3+sllzPWVL63pAJdwAircb4FzJYsRVWX9DZ/cRJhbsPUjngXu1E1ZYWhDPw6VTtk+V1yMXap+65xvc23pjl0MUl5m/jDcmdh9C6Dr1dtnGm1k17xSD1WOGW9QWNFpM42jZdaZhNKD+Tcf/5wxZBaEzhjIzMbInE/OhBIBFb837e1lbrrE44oSigxwTIUWDi6TBqlCLvUCSOe+N6O65+gS4kIEoFGZzcUEnbMPAI6TFbQEbhfzcvUjHV1IXKWR5VSd2Chkhd7sK61npWrSAbX2YTZ09TDdwotKt1LBuyU+A7ObJR24alubeLcnh4NOiJPPdQBnX69U3HOwa8OtnnUsMhuz/V6o2tXCnTwtPFZOu0ofu2ZDIN6/xpH/uCF3SEBba9eXOw==
x-forefront-antispam-report: CIP:255.255.255.255;CTRY:;LANG:en;SCL:1;SRV:;IPV:NLI;SFV:NSPM;H:SJ1PR12MB6289.namprd12.prod.outlook.com;PTR:;CAT:NONE;SFS:(13230040)(376014)(366016)(1800799024)(38070700018);DIR:OUT;SFP:1101;
x-ms-exchange-antispam-messagedata-chunkcount: 1
x-ms-exchange-antispam-messagedata-0: B7nf0j1vUXaBgqxENcIFvS6BAQriTUNoIar9fa+gI4PxgWdSwQyykGWM0cIwHkPdsiEhiiOwP/OLMMRFA7WZ9raMuWfKpngO6a6Iw0RTPihVN2+hkchSFvBjuEqYnmVgj1l9OODJaddLEItASv4bT+6TwWTfnsA4YaDO2cCqBB7l3cmZqc1rVOX8TkXde2xZkXCYqzw5lpV7NAZLFzxDN/7mpJAKBuGq3pfXqAN5Oj9MfAJ0K65OUrH5giAqMUkFFeQAzhG/+vF8PAZWTmC07kVoTGc5Z17WTpBiQ98FJ+tafqdRlCg3x4bvJlbRR0CsKZOwSabOewr+VfKS9TPPSCmHxKtFpWkrXhq4bwTKNwINjyd5DxqkoYKMNV1XSTIazSwiydzD/r60X/BCnlsb17LpC4+Fdfg8M83R+yptHeV41uCNIqfg4T0iWaoUH85j6DC8C0a2M4Os8UWTUqAMj1BGntSL6Juxlca8enBF6x0CBz6JAkd2WE4KSCfMEom+bjCVMgWNfcJ5Xuu9d0TqnG/pmMDthtCtgv9gj4evjXvb55p3C3mZ6lGZlaefCetfq8qNf9oB26r5PUKoQMvIzOUZlssyuvci0p0LIGZM+3IpnTkdj/GaUko6tY7kegHpBaP8T3AdVFlBWDrVq8DwOhbPMDkwvx5qV575tO22KTt1gu+PkaqPbd9iO+RRSbPDfwPmw7G8rOHEIvVgpB+Uu/2hbc3NUJboVmv+X5jswNWDscNoxj07Qh6fsblRfmIJ8un8Ure2lgzcEYnbnpFAClsQZekK/tmbVeCR8UYQSsEGy3hcOHtuUt0r0OEz4TegATJ7NXJfrS8ArviVITjr9WQVz7Yq40i0WrYPHipyUYvoCl4mLy1e636lADObluQa7sKfK6JdSTq+PojcHHJXs2mq8z9szXp0xkZGulaKizKRm104/W2CJG+KimRNXyXDbf8ea1qAywv84o8Kji7jA3w/Zf+an9buYysWpBmgmw5vNUGa5wkI5rhNH6d5alguyW2THx20JwT/Dz8LSlY58dm+rrUgUREuEN9H1lp44Iokcaew5pjbaRNzSVLqfR7K60uKVfgQNeLFR93p1Z9hIu6VVQ24T/B3x6JQ1U+RZHhMEB91YofUsLpCqjQykpqQPHIrm5ia5wsHj/+OlDeGh6c2/DpwLHXtc1Ob7eJn6uCIq8v+7vVAYjllRAWRYiKhlALSZIsTt3cTyJnwySAZaKsb6tPoD8nCQ8pVgFy0wVVwTJpGdAeywNaxK4vh3Vfi6t5LQj/udT96zbr2VvSsFoLoKfU/lN1xEMq1WRhJflFKb7NaKAPftRs5B3WHvhhupvolt6UVs5Io800e3XCAAwnOBHSYDFKsfsuyga54JMsJ8N2+N6WnFytk19r8A9XCSsOOT/heqRAMyVewdpGCZmOVrRZD2jz6Jk4q1I8m6LsLFEBP0VWMNHdDqWU/uTweTZ98FT6X/fZEeArAH667YJcIHmEo79AuDedzsM88DE9LJXa++IT5vFyj/LHoW/lTROpij//qvM9TlpjeTfWy3UOduA3QsVpvLT19U2Y+O4x49FPeSNfoz4lljgJrZQNw
Content-Type: text/plain; charset="us-ascii"
Content-Transfer-Encoding: quoted-printable
MIME-Version: 1.0
X-OriginatorOrg: Nvidia.com
X-MS-Exchange-CrossTenant-AuthAs: Internal
X-MS-Exchange-CrossTenant-AuthSource: SJ1PR12MB6289.namprd12.prod.outlook.com
X-MS-Exchange-CrossTenant-Network-Message-Id: 7ca2727c-dc4c-4338-e7e6-08dcd743e654
X-MS-Exchange-CrossTenant-originalarrivaltime: 17 Sep 2024 18:09:40.6759 (UTC)
X-MS-Exchange-CrossTenant-fromentityheader: Hosted
X-MS-Exchange-CrossTenant-id: 43083d15-7273-40c1-b7db-39efd9ccc17a
X-MS-Exchange-CrossTenant-mailboxtype: HOSTED
X-MS-Exchange-CrossTenant-userprincipalname: f0rAmzU8mPI4c2IDMpXKucFlk0YoYhiNsDRupN+pLlKIebZ6nOrNUNG8pBzOOkzerw3ncssoRai09XYEIKLQXg==
X-MS-Exchange-Transport-CrossTenantHeadersStamped: DS0PR12MB7969
X-MailFrom: vkushwaha@nvidia.com
X-Mailman-Rule-Hits: nonmember-moderation
X-Mailman-Rule-Misses: dmarc-mitigation; no-senders; approved; emergency; loop; banned-address; member-moderation; header-match-art.ietf.org-0
Message-ID-Hash: A736MZZEAW7KLCVA2DQCRRUDBQN6HUFH
X-Message-ID-Hash: A736MZZEAW7KLCVA2DQCRRUDBQN6HUFH
X-Mailman-Approved-At: Fri, 08 Nov 2024 03:27:03 -0800
X-Mailman-Version: 3.3.9rc6
Precedence: list
Subject: [art] Re: Gdeflate as RFC
List-Id: Applications and Real-Time Area Discussion <art.ietf.org>
Archived-At: <https://mailarchive.ietf.org/arch/msg/art/z_F7ooIC68517jTkmCiL7uCpJAc>
List-Archive: <https://mailarchive.ietf.org/arch/browse/art>
List-Help: <mailto:art-request@ietf.org?subject=help>
List-Owner: <mailto:art-owner@ietf.org>
List-Post: <mailto:art@ietf.org>
List-Subscribe: <mailto:art-join@ietf.org>
List-Unsubscribe: <mailto:art-leave@ietf.org>
Date: Tue, 17 Sep 2024 18:09:49 -0000
X-Original-Date: Tue, 17 Sep 2024 18:09:40 +0000

Hi John,

We are not suggesting this as a replacement for Deflate or for use in standard systems. Instead, we aim to propose it as an RFC, allowing implementers with parallelism-friendly hardware to reference and adopt it in their systems if they choose.

~Vikram

-----Original Message-----
From: John R Levine <johnl@taugh.com>
Sent: Tuesday, September 17, 2024 2:03 PM
To: Vikram Kushwaha <vkushwaha@nvidia.com>; Akshay Subramaniam <asubramaniam@nvidia.com>; Martin Thomson <mt@lowentropy.net>; art@ietf.org
Subject: RE: [art] Re: Gdeflate as RFC

Thanks for your note.  This is all informative but it still doesn't give me an understanding of why we would want to add this algorithm as an additional one for HTTP.

It's certainly faster in some circumstances but I'm not seeing how it'd offer an interesting difference on normal systems like PCs or phones that have a few CPU cores and a few GPU cores.

R's,
John

On Tue, 17 Sep 2024, Vikram Kushwaha wrote:

> Hi Martin/John,
> Please let us know if you have any more questions on GDeflate performance.
>
> Thanks!
>
> ~Vikram
>
> From: Akshay Subramaniam <asubramaniam@nvidia.com>
> Sent: Thursday, August 29, 2024 6:58 PM
> To: Vikram Kushwaha <vkushwaha@nvidia.com>; Martin Thomson
> <mt@lowentropy.net>; John R Levine <johnl@taugh.com>; art@ietf.org
> Subject: Re: [art] Re: Gdeflate as RFC
>
> Hi all,
>
> I can't read the full history of this email thread, especially plots or other attachments that were shared but let me add some thoughts on GDeflate vs ZSTD.
>
> I'm assuming the plots that you saw were some fairly coarse grained average compression ratio and throughput plots. Some clarifications on that:
>
>  1.  Both GDeflate and ZSTD in those plots are implemented on the GPU. Specifically, the ZSTD implementation used there is a custom CUDA based GPU implementation in the nvCOMP library, not the public CPU ZSTD library.
>
>  1.  The compression ratios for GDeflate are higher than ZSTD mainly because the compression techniques are different. Since GDeflate was designed for fast decompression, we implemented a high compression version of the compressor (similar to libdeflate level 12). Our GPU implementation of ZSTD was developed mainly for database applications where both compression and decompression throughput matter and so we tradeoff some compression ratio for throughput. The difference in compression ratio is only a result of the compressor implementations, not of the stream formats themselves.
>
>  1.  The speed difference between GDeflate and ZSTD depends a lot on the dataset. The main innovation in GDeflate is to make the entropy decoding much faster with the swizzling technique. But if the decompression performance is bottlenecked by LZ decompression, then the gap between GDeflate and ZSTD would be smaller.
>
>     *   Here's an example just for the Silesia corpus, GDeflate decompression throughput is 53 GB/s while ZSTD (on GPU) is 39 GB/s. So GDeflate is ~36% faster in this case. If we take a dataset that is much more entropy decode bound, the difference might be expected to grow.
> I want to reiterate that the ZSTD numbers are from a highly optimized GPU implementation in nvCOMP. If you look at the performance of the regular ZSTD libary on lzbench<https://github.com/inikep/lzbench>, it tops out at 1.2GB/s.
>
> There is one other use case for GDeflate that might be interesting. Since the main innovation in GDeflate is from faster entropy coding, we can use a very fast throughput compression algorithm by only entropy coding data without the LZ phase. This allows for symmetric compression and decompression throughputs of 150-200 GB/s and allows for compression to be used in communication bound applications where communication is done over relatively high performance networks.
>
> Thanks,
> Akshay
> ________________________________
> From: Vikram Kushwaha
> <vkushwaha@nvidia.com<mailto:vkushwaha@nvidia.com>>
> Sent: Thursday, August 29, 2024 10:36 AM
> To: Martin Thomson <mt@lowentropy.net<mailto:mt@lowentropy.net>>; John
> R Levine <johnl@taugh.com<mailto:johnl@taugh.com>>;
> art@ietf.org<mailto:art@ietf.org> <art@ietf.org<mailto:art@ietf.org>>;
> Akshay Subramaniam
> <asubramaniam@nvidia.com<mailto:asubramaniam@nvidia.com>>
> Subject: RE: [art] Re: Gdeflate as RFC
>
> I am not familiar with ZSTD, but adding @Akshay Subramaniam who will be able to better answer questions on the ZSTD vs Deflate performance.
>
>
> ~Vikram
>
> -----Original Message-----
> From: Martin Thomson <mt@lowentropy.net<mailto:mt@lowentropy.net>>
> Sent: Wednesday, August 28, 2024 8:52 PM
> To: Vikram Kushwaha
> <vkushwaha@nvidia.com<mailto:vkushwaha@nvidia.com>>; John R Levine
> <johnl@taugh.com<mailto:johnl@taugh.com>>;
> art@ietf.org<mailto:art@ietf.org>
> Subject: Re: [art] Re: Gdeflate as RFC
>
> Thanks for sharing that Vikram, it's helpful.
>
> I'm curious as to what zstd folks think about these results.  Is it the case that the tuning of zstd was to ensure that the throughput would be comparable to gdeflate?  From my understanding, zstd is capable of far better compression ratios than deflate and that your work was primarily focused on throughput, such that gdeflate would necessarily outperform zstd on that axis as much as it seems to have done in this scenario.
>
> I'm not sufficiently expert here, so I'll defer to others.  However, those numbers aren't necessarily convincing (though the OOMs on zstd might be if it were explained).  zstd seems to have slightly less throughput and compression both, but it's not so clearly a win as I'd have expected.  LZ4 is clearly unsuitable, but that's expected.
>
> On Thu, Aug 29, 2024, at 06:29, Vikram Kushwaha wrote:
>> Hi John/Martin,
>>
>> As I continue to work on the IPR disclosures, as requested, attaching
>> comparisons of GDeflate performance with other compression algorithms.
>> The summary is that GDeflate has a high throughput while maintaining
>> a high compression ratio.
>>
>> Thanks,
>> ~Vikram
>>
>> -----Original Message-----
>> From: Vikram Kushwaha
>> <vkushwaha@nvidia.com<mailto:vkushwaha@nvidia.com>>
>> Sent: Monday, August 5, 2024 9:51 PM
>> To: John R Levine <johnl@taugh.com<mailto:johnl@taugh.com>>; Martin
>> Thomson <mt@lowentropy.net<mailto:mt@lowentropy.net>>;
>> art@ietf.org<mailto:art@ietf.org>
>> Subject: RE: [art] Re: Gdeflate as RFC
>>
>> Thanks all for your feedback.
>>
>> I am still reading up on the IP disclosures, I will follow up on that
>> in an another email.
>>
>>> What sort of applications do you have in mind here?
>> This will be useful in applications where GPU decompression can be
>> done on the fly, offloading the work from CPU. One use case is gaming
>> where frames are streamed and data asset decompressions are handled
>> by some of the GPU cores.
>>
>>
>>> If it's supposed to be generally useful it'd also be helpful to have some idea how it works on normal CPUs.
>> That would depend on how the decompressor is written. If a CPU based
>> decompressor can make use of the parallel nature of the decompression
>> with gdeflate(it should be with thread programming) it will be
>> faster, though we haven't written one and so do not have numbers for
>> comparison as CPU cores vary a lot.
>>
>> I am going to send out an email comparing gdeflate with zstd and other formats.
>>
>> Of course, lower end GPUs will show a smaller gain but it will still
>> be faster than CPU decompression. The advantage we are going for is
>> by offloading decompression to GPU, CPU can be used for other tasks.
>> This is quite useful in gaming/visualization apps where CPU to GPU
>> communication is a bottleneck.
>>
>> One more thing I would like to add is that Microsoft has already
>> adopted GDeflate as their default GPU decompression method, so by
>> promoting this as a RFC we were hoping to standardize it more globally.
>>
>> ~Vikram
>>
>> -----Original Message-----
>> From: John R Levine <johnl@taugh.com<mailto:johnl@taugh.com>>
>> Sent: Thursday, August 1, 2024 8:29 PM
>> To: Martin Thomson <mt@lowentropy.net<mailto:mt@lowentropy.net>>;
>> art@ietf.org<mailto:art@ietf.org>
>> Cc: Vikram Kushwaha
>> <vkushwaha@nvidia.com<mailto:vkushwaha@nvidia.com>>
>> Subject: Re: [art] Re: Gdeflate as RFC
>>
>> On Fri, 2 Aug 2024, Martin Thomson wrote:
>>> On Fri, Aug 2, 2024, at 06:38, John Levine wrote:
>>>> What sort of applications do you have in mind here? The main place
>>>> DEFLAATE is used in IETF protocols is in compressed HTTP streams
>>>> which doesn't strike me as the kind of thing one would often do on a GPU.
>>>
>>> I don't see why not, if the GPU is that much faster.
>>
>> That benchmark is on an RTX3090 which has 10,000 cores.  I'm typing
>> this on an M2 Pro laptop whose GPU has 19 cores.  I believe their
>> benchmark but I also don't see how it's relevant to anything but high
>> end dedicated GPUs.
>>
>>> I'm more concerned about the comparison to more modern stuff, like
>>> brotli or zstd.  Those tend to offer far better compression
>>> performance than deflate. ...
>>
>> That is an excellent point if we're going to change the comprsssion scheme.
>>
>> R's,
>> John
>>
>>
>> Attachments:
>> * GDeflateComparison.pdf
>

Regards,
John Levine, johnl@taugh.com, Taughannock Networks, Trumansburg NY Please consider the environment before reading this e-mail. https://jl.ly/