Re: [AVTCORE] [External] Re: Comments on “RTP Payload Format for Versatile Video Coding (VVC)”

"Hannuksela, Miska (Nokia - FI/Tampere)" <miska.hannuksela@nokia.com> Mon, 30 August 2021 18:38 UTC

Return-Path: <miska.hannuksela@nokia.com>
X-Original-To: avt@ietfa.amsl.com
Delivered-To: avt@ietfa.amsl.com
Received: from localhost (localhost [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 718E93A1D55 for <avt@ietfa.amsl.com>; Mon, 30 Aug 2021 11:38:00 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -2.353
X-Spam-Level:
X-Spam-Status: No, score=-2.353 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIMWL_WL_HIGH=-0.452, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, HTML_MESSAGE=0.001, RCVD_IN_MSPIKE_H2=-0.001, SPF_PASS=-0.001] autolearn=ham autolearn_force=no
Authentication-Results: ietfa.amsl.com (amavisd-new); dkim=pass (1024-bit key) header.d=nokia.onmicrosoft.com
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id GC_FFvF_2qD5 for <avt@ietfa.amsl.com>; Mon, 30 Aug 2021 11:37:55 -0700 (PDT)
Received: from EUR05-DB8-obe.outbound.protection.outlook.com (mail-db8eur05on2106.outbound.protection.outlook.com [40.107.20.106]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by ietfa.amsl.com (Postfix) with ESMTPS id BC6993A1D51 for <avt@ietf.org>; Mon, 30 Aug 2021 11:37:53 -0700 (PDT)
ARC-Seal: i=1; a=rsa-sha256; s=arcselector9901; d=microsoft.com; cv=none; b=nw8Hz1Lg1tetkM3BimWXrwQWFIc8i7CRcVKJ/c+K1OFPmGjghyp+mLuhrE+4NSaR7iEzkJCWOqccyNTnW7Mf7PGe4wKWAzLnP5Hv1hW7dqwFQvaLcURFe8HZC3xR+9fkpFeZn2yNWRjHrqoK16zAfBLfYeZ4PIYY6c9AjG3u1ZgQWTDnWbaIOrsrqP8kaoWa33SYgyQkifKeiGEb/yDDTXlyviPAOMIsB+Xk7E4FrjYZRrnMIP5zCXygcHb+OHBiH6QQvwZJWgvhCZPZWoEbqiQu6K2ZUbyRzWBEvkV9U9u4+SCEco91fCGqXqdLhN/2dgRhYwhToV/6DGFDBMANiA==
ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector9901; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version; bh=QObEm0jt0KT8Cvum96p2M1cr8UkmeVH28RMQm1t1ZHE=; b=dlHpcmtqmoJ4RlETFqsg0dmDN5sMhU4lcS5H4wL7BoLQt8dzCRbDJYpt5ZBj7oKxhQaXM/J4Br1iaW+65wokBcwA6r9DrnVs+8Fw9kVQ8sf5/F1dK0bAIPFTFLVKgoYmEC2AaAFiEUVUoUTmxs83rS8VRwTpAVeWH68h1UvgBHd0eBxf4lM7ggX3Kw9YF9m5OzJ5uChxnz6Qf8L91AjpLHNd8qq0pw1op9d77x+vS8QLPvCQlqtbYxk+qt2jRa4Lw4YV0lIuhjX7usA1GGjr9pjiTAEYdGz9sYT7Fw/ggnNVnFVsYX6xPfVa2hp/8rtA6kS6fIBzxvpGucu/VuNBEw==
ARC-Authentication-Results: i=1; mx.microsoft.com 1; spf=pass smtp.mailfrom=nokia.com; dmarc=pass action=none header.from=nokia.com; dkim=pass header.d=nokia.com; arc=none
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=nokia.onmicrosoft.com; s=selector1-nokia-onmicrosoft-com; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-SenderADCheck; bh=QObEm0jt0KT8Cvum96p2M1cr8UkmeVH28RMQm1t1ZHE=; b=NbKEhZ512opEnMkjG3b0y9B/qSG2qc7dyuW3aTskpYynqCdItxPw9Np4wIhtVHDK0kobiA+qMgXBmWS0aYBuSwtQIEAWyS92lzbfC2u7l/yOCCz6FBZ8Hq5eO1hb38QGkOJ8+9arcXu2Fn2m+hg8J4NZt3ASyXQIDMG91g3lM1Y=
Received: from HE1PR0701MB2857.eurprd07.prod.outlook.com (2603:10a6:3:57::7) by HE1PR0701MB2891.eurprd07.prod.outlook.com (2603:10a6:3:4b::12) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.20.4478.10; Mon, 30 Aug 2021 18:37:46 +0000
Received: from HE1PR0701MB2857.eurprd07.prod.outlook.com ([fe80::4169:6369:8d9:2cc6]) by HE1PR0701MB2857.eurprd07.prod.outlook.com ([fe80::4169:6369:8d9:2cc6%5]) with mapi id 15.20.4478.017; Mon, 30 Aug 2021 18:37:46 +0000
From: "Hannuksela, Miska (Nokia - FI/Tampere)" <miska.hannuksela@nokia.com>
To: "Sanchez de la Fuente, Yago" <yago.sanchez@hhi.fraunhofer.de>, Ye-Kui Wang <yekui.wang@bytedance.com>, Stephan Wenger <stewe@stewe.org>
CC: IETF AVTCore WG <avt@ietf.org>, "shuaiizhao(Shuai Zhao)" <shuaiizhao@tencent.com>
Thread-Topic: [External] Re: [AVTCORE] Comments on “RTP Payload Format for Versatile Video Coding (VVC)”
Thread-Index: AQHXmfi17kQsMqMvO0G1N3wxltqHnquGaUMAgAD8HwCABOc6gA==
Date: Mon, 30 Aug 2021 18:37:46 +0000
Message-ID: <HE1PR0701MB2857A8342F7CD1BC0D9E209F88CB9@HE1PR0701MB2857.eurprd07.prod.outlook.com>
References: <78FB6245-8BA4-4C47-8C00-CCC6D170A68F@stewe.org> <042f01d79ace$9bdd48b0$d397da10$@bytedance.com> <16DEF487-5725-4D6D-8CD1-39CB533A3FA1@hhi.fraunhofer.de>
In-Reply-To: <16DEF487-5725-4D6D-8CD1-39CB533A3FA1@hhi.fraunhofer.de>
Accept-Language: en-US
Content-Language: en-US
X-MS-Has-Attach:
X-MS-TNEF-Correlator:
authentication-results: hhi.fraunhofer.de; dkim=none (message not signed) header.d=none;hhi.fraunhofer.de; dmarc=none action=none header.from=nokia.com;
x-ms-publictraffictype: Email
x-ms-office365-filtering-correlation-id: 1a57f5ab-5ce0-4006-f56b-08d96be542ee
x-ms-traffictypediagnostic: HE1PR0701MB2891:
x-microsoft-antispam-prvs: <HE1PR0701MB2891D8E87CE23B473FF79CFB88CB9@HE1PR0701MB2891.eurprd07.prod.outlook.com>
x-ms-oob-tlc-oobclassifiers: OLM:10000;
x-ms-exchange-senderadcheck: 1
x-ms-exchange-antispam-relay: 0
x-microsoft-antispam: BCL:0;
x-microsoft-antispam-message-info: X5htxoTbFIoTxdu8buETyzifVEavTgkgAILwx8MXY5kXjbjFtqPtnlNx9WwKh1qLDpYfWAOzp+TuIoWUpWy1zNBd9VX+GjW47+505FM+Sb2tFBdEuCyaUpLV4UCH5Mv6LjZ4vPlqXIq6XnnkcNIYISPRKD1b/idkJ4MHTaSLZLUxiSl5UNzieTsIQJ0VrP283ZNgVcwujiyom4JrhuybrJlipk1i6DaxCjXUvXsa0IQeVtZYdAzvp2SkPcmNtu34IzP30a3W25XQfqYeqfTXtL5kg1YcwOfKU5Pq7ZGvsF9zK9l0EVezzn8M8RftLA2NMbWJsn8zZ7J6DxkaBG8VknpQ0TDOaUWnDfs68JmnZ0DCJJo2PFGauIrPXL8XZTtkUlIcnXCT3hbiu08+EEEAE8h+NhRrHG6zHQP1VS95wKnmDCRefFHNfT6kaqrORXpwLniwlkErmKAO2DBYVBV0dF3sJVF9N6KQTIzNdu2Z4uNJEVDZtkGgNSHQJ5oUttSiKEExcdV4YWX70Emg2+dasm+2U7uDMT0lW1Tv0STPlTCc0H4I5FKkeji9tDs0PFOyp+FCxub9HprLE94FYExcSSy0bpMInt7FRoPxDYAIfIH7FZWyoAeRyJxp9xV8NneLzZx8QAUtD6wOAiM43GujYnszIJwTq/PrVT98NGpgPU6OOb+6oMboRllJDoDfnFMnPCuk53YnM9spqwfV7q64OQ==
x-forefront-antispam-report: CIP:255.255.255.255; CTRY:; LANG:en; SCL:1; SRV:; IPV:NLI; SFV:NSPM; H:HE1PR0701MB2857.eurprd07.prod.outlook.com; PTR:; CAT:NONE; SFS:(4636009)(366004)(66946007)(30864003)(26005)(66446008)(66476007)(64756008)(66556008)(8936002)(5660300002)(33656002)(316002)(2906002)(110136005)(54906003)(4326008)(508600001)(38070700005)(55016002)(52536014)(76116006)(6506007)(83380400001)(38100700002)(122000001)(186003)(7696005)(71200400001)(86362001)(9686003)(579004)(559001); DIR:OUT; SFP:1102;
x-ms-exchange-antispam-messagedata-chunkcount: 1
x-ms-exchange-antispam-messagedata-0: fQZugEeMetlsdsdC9leqHtwmOUQ+xqPbOsswi6/RUYfLdQwSvbqA63eW82L0/RGzRqlTWkFAIi/IH6wALTT8J82gnH6WaaPySlNbEMq/wEXuXMlqxZQ2K4VXLTIhrmZ/EjeaFm/QDSlIZZKi7QibKAayEs92SWDzZXOcHxNRlRfkUpOIgfoO05pjSvzt6ficto0LwOjyVR995bucN/WK0wWySQVxxKW1DbpExdpDi7r3Tpj32s+btDr2zpBPiDw0kvlAAIjio0Jff1HIPRCEkok8c1aMMHnLAQL0tAeflaSmVhfGbuUpnC0LLCc8LaPrVn6MfqyDtZBftJYgqf1JSmH0hEfx+tE00YnDYO3Zd71Rgfcz8JiJThxi3ZabPtHtlbePaW97F5zhp5Xmuf6zt6gefSCJzWDX+VaWxNb+nATsNgOwaxcR1de0XytABQ/JyJluZCT86hDgwv6m7tbOI684cTMBlRZ0eytF9oFIU9igzXkFctJxIUiEtmlou2RyyeeskBQHqOQMICkjBBliranLFt6cIudUkQ4BylQ6oevX4N8/2GYBx+qzg64gsjleig97JL7XbQZzZjHLkIeQrwqb+B+m9RBTq9j2anLBHNKrc1YLrpMpmC90G0HxSAipa+Ft35kZ0fJFlkq12qzl0m/xY5yPK9NnH85L9lAiAtfFIHod0BTgrc0Ig2E9UmsHwxv163RAdxLsjnYc5vAirbB1nlO6aauL8Pl8uu0sanYjJgMTujHetkADlz/w5ZfUHOCtF7n2J6GF4VRGD35zzpi1Wokyk02bVrutNQLIz1nL3QiXX9yYcc1H2qOc8euvqKfpydBrnEx5h0rUim9cC7y3FLNDuEOlUZKkzU8CHYAwLjpalZ+X+waVK1kQbKGwVo6d5kR2ArQxjPnHGd3PYuzhgQtWKLz0nexS4+kf4AMvFdc021ojfnzQaAJJlO6WSCvrpxR/LkCP55UHsMhjGbJCdF7phf+f6d1tqort8fSG8Foo9/Ahb206kA8GLFzD0D7ESShm6YrWc/64lulVc9Zc9RfR8i3XvJMGpy/BxhDjiuEFpfMPaR3ldOIKZXn2MCJs4l7b9TLrGOjz5tkf/kdqp051cNJDoPmbWvKRYXjlz4mQ5cVYdApcR93tzDOaDWEVtCL1w2WAv3KBW4dCqVMeLQlHk4qLVB58neemRDMH3da4zkD3MoqPqPtDLnJHKGIGy2T1n2boK+/1xFbe6bB7Yp72tAW3iWSf89QZxhGvJQNl6GMjky9pB+s+83pKhGN8Wx4JkXvpiOuGLRi9Vwuq4ctUm3G4GLH86HMdzNCUbFC3a1NeEQsq/R+g8UIW
x-ms-exchange-transport-forked: True
Content-Type: multipart/alternative; boundary="_000_HE1PR0701MB2857A8342F7CD1BC0D9E209F88CB9HE1PR0701MB2857_"
MIME-Version: 1.0
X-OriginatorOrg: nokia.com
X-MS-Exchange-CrossTenant-AuthAs: Internal
X-MS-Exchange-CrossTenant-AuthSource: HE1PR0701MB2857.eurprd07.prod.outlook.com
X-MS-Exchange-CrossTenant-Network-Message-Id: 1a57f5ab-5ce0-4006-f56b-08d96be542ee
X-MS-Exchange-CrossTenant-originalarrivaltime: 30 Aug 2021 18:37:46.2056 (UTC)
X-MS-Exchange-CrossTenant-fromentityheader: Hosted
X-MS-Exchange-CrossTenant-id: 5d471751-9675-428d-917b-70f44f9630b0
X-MS-Exchange-CrossTenant-mailboxtype: HOSTED
X-MS-Exchange-CrossTenant-userprincipalname: izN04BVNDTAphjjlKHqkJRwfRyyc5BmZgS6ed9i23aYCdXMQ7l26ySPEdZ9c6bPUSOIVOoO50urYaf4kETHpigQkhgRPmHG0Jq2hdQzTJ6E=
X-MS-Exchange-Transport-CrossTenantHeadersStamped: HE1PR0701MB2891
Archived-At: <https://mailarchive.ietf.org/arch/msg/avt/RZ9J-y_iAZbqnb701bdKhvDf1wk>
X-Mailman-Approved-At: Tue, 31 Aug 2021 09:13:24 -0700
Subject: Re: [AVTCORE] [External] Re: Comments on “RTP Payload Format for Versatile Video Coding (VVC)”
X-BeenThere: avt@ietf.org
X-Mailman-Version: 2.1.29
Precedence: list
List-Id: Audio/Video Transport Core Maintenance <avt.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/avt>, <mailto:avt-request@ietf.org?subject=unsubscribe>
List-Archive: <https://mailarchive.ietf.org/arch/browse/avt/>
List-Post: <mailto:avt@ietf.org>
List-Help: <mailto:avt-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/avt>, <mailto:avt-request@ietf.org?subject=subscribe>
X-List-Received-Date: Mon, 30 Aug 2021 18:38:01 -0000

Hi Stephan, Ye-Kui, Yago,

Thanks for your responses. Below I am continuing the discussion on items that have not been concluded yet.

2) Section 7.1, profile-id semantics

         A profile_tier_level( ) syntax structure may be contained in an
         SPS, VPS, or DCI NAL units as specified in [VVC].  One of the
         following three cases applies to the container NAL unit of the
         profile_tier_level( ) syntax structure containing those PTL
         syntax elements used to derive the values of profile-id, tier-
         flag, level-id, sub-profile-id, or interop-constraints:
         ...
         3) The container NAL unit is a DCI NAL unit and the
         profile_tier_level( ) syntax structures in all DCI NAL units in
         the bitstream has the same values respectively for those PTL
         syntax elements.


  *   The first sentence above might give a wrong impression, since VPS and DCI may contain multiple PTL syntax structures. The sentence could be rephrased as follows: “As specified in [VVC], a profile_tier_level( ) syntax structure may be contained in an SPS NAL unit, and one or more profile_tier_level( ) syntax structures may be contained in a VPS NAL unit and in a DCI NAL unit.”

It's of course correct that any of the parameter sets mentioned can include more than one PTL structure, and we will adjust the text accordingly.
I don’t recall why multiple PTLs were allowed, but I don’t think it’s critical for an RTP-based system, where real-time encoders are the norm and splicing the exception.  Perhaps the easiest way to deal with problems arising from potentially multiple PTL structures (as you outline below), is to discourage those bitstreams.  Something like “VVC bitstreams transported over RTP using the technologies of this memo SHOULD contain only a single PTL structure in the DCI, VPS, and SPS.  The VVC syntax is more flexible, in allowing multiple PTL structures, but some of the mechanisms described herein would show undefined behavior.  If a system were to employ multiple PTL structures contrary to above “SHOULD”, the encoder(s) need to ensure that the mechanisms of this payload format, which are designed for only a single PTL structure, continue to work, for example by populating the multiple PTL structures with non-contradictory information.”

[YS] I think that Stephan’s suggestion here is good while I think that it only is needed for DCI. We might have several PTLs in VPS for different OLSs and that should be fine.

[MH] I agree with Yago that discouraging the use of multiple PTLs only concerns DCI NAL units.



  *   I am a bit uncertain how to interpret the phrasing “the PTL syntax structures in all DCI NAL units in the bitstream has the same values respectively for those PTL syntax structures”.

     *   Let’s take an example: there are two DCI NAL units, dciA and dciB, in the bitstream, and both of them contain two PTL structures, i.e. dciA contains ptlA1 and ptlA2 structures (in this order) and dciB contains ptlB1 and ptlB2 structures (in this order).
     *   Does the phrasing mandate ptlA1 to be the same as ptlB1, and ptlA2 to be the same as ptlB2? If so, then why is the phrasing needed at all, since the VVC standard requires all DCI NAL units in a bitstream to have the same content.
     *   Or does the phrasing mandate ptlA1 to be the same as ptlA2, and ptlB1 to be the same as ptlB2? This interpretation would not be meaningful technically, since why would an encoder include multiple PTL structures in a DCI NAL unit, if they the same values.

I think this would be solved through the above limitation.

  *   It seems that case 3 (DCI) would not work if a DCI NAL unit contains more than one PTL syntax structure. If there were multiple PTL structures in a DCI NAL unit, how would it be determined which one of them is used to derive profile-id, tier- flag, level-id, sub-profile-id, and interop-constraints? Would case 3 need to be limited to apply only if DCI NAL unit(s) contain only one PTL structure?

Yes, to the second question, and again, the problem would be solved through above limitation.


[YS] I think this is ok, but I am unsure what would we say when there is more than on PTL. What would we exactly mean with “for example by populating the multiple PTL structures with non-contradictory information”? Would it  mean that all PTLs are the same? In that case we are good. But if the meaning of non-contradictory PTLs is more relaxed, we might need to explain here a bit more. Let’s for instance take as an example two PTLs in a DCI with each indicating a different sub-profile. I would understand non-contradictory information that one of these profiles needs to be a subset of the other and not simply “disjoint” profiles. Then we would need to say that the sub-profile in the SDP needs to be the one that includes the other sub-profile. Unless we completely disallow different PTLs in a DCI, in which case the current text is ok since both sub-profiles would be exactly the same. If the later, we should clearly indicate that two PTLs in DCI are not supported but the “SHOULD” above would not be enough I think. Which one should we do?

[MH] Note that this part of the I-D discusses how the sender can derive the value for profile-id from 1) SPS, 2) VPS, or 3) DCI. Using case 1 or 2 for deriving profile-id is always possible, while DCI with a single PTL can simplify the operation. I’d tend to think that allowing multiple PTLs in DCI NAL unit is still allowed but discouraged (as Stephan suggests above). If a DCI NAL unit contains multiple PTLs, the sender needs to derive the profile-id from SPS or VPS.


5) Section 7.1, recv-ols-id

         When present, the value of recv-
         ols-id must be included only when sprop-ols-id was received and
         must refer to an output layer set in the VPS that is in the
         same dependency tree as the OLS referred to by sprop-ols-id.


  *   I don’t understand why sprop-ols-id is required in the offer in order to have recv-ols-id in the answer. Could you explain?

     *   Note that if this requirement is relaxed, Section 7.2.2.2 would also need changes.


  *   In my understanding:

     *   The absence of sprop-ols-id in an offer indicates that the offerer can transmit any OLS that is included in sprop-vps. Consequently, the answerer could respond with recv-ols-id equal to any OLS ID included in sprop-vps.
     *   The presence of sprop-ols-id in an offer indicates that the offerer only has a bitstream corresponding to the OLS indicated by sprop-ols-id and can transmit only that OLS or any OLS contains either the same layers as or a subset of the layers of the OLS indicated by sprop-ols-id.

Yago, can you take this?  I vaguely recall discussions around this, but not the conclusion.

[YS] We currently have a default value for sprop-ols-id in section 7.1, which follows the same inference rule as TargetOlsIdx in 8.1.1 in [VVC], so this means to me that the absence does not necessarily indicate that any OLS included in the sprop-vps can be transmitted. At least this was not my recollection. If we wanted to do so we should remove the inference value and state this clearly.  As for why we took this design choice, I remember we had some discussion about OLS negotiation and our intention was to keep things simple. We wanted to simplify the OLS negotiation and we decided that 1) the presence of sprop-ols-id was indicating the willingness of the sender to negotiate OLSs and the absence of it meant that that flexibility was not desired and “regular” negotiation should be carried out where the media format is included as is in the answer or removed completely (when one or more of the parameter values are not supported); and 2) when negotiating OLSs not anyone could be selected by the answerer but only a subset of the offered one (the text regarding the dependency tree).

[MH] Thanks for the explanation. With Stephan’s definition of dependency tree (below), I am able to understand all the pieces now. I don’t have strong opinions on this issue. If design stability is preferred, the current text can be kept except that I find the inference of sprop-ols-id confusing, since the inferred value is not used for anything as far as I can see. Therefore I would propose to delete the sentence on inferring the sprop-ols-id value. Or am I missing something?



  *   The term dependency tree is not defined and hence should either be avoided in the phrasing or defined.

Does this work: “[…]must refer to an OLS in the VPS that includes no layers other than all or a subset of the layers of the OLS referred to by sprop-ols-id. “

[MH] Sounds good, thanks.


6) Section 7.1, sprop-vps

         The sprop-vps parameter MAY contain one or more than one video
         parameter set NAL unit.  However, all other video parameter
         sets contained in the sprop-vps parameter MUST be consistent
         with the first video parameter set in the sprop-vps parameter.
         A video parameter set vpsB is said to be consistent with
         another video parameter set vpsA if any decoder that conforms
         to the profile, tier, level, and constraints indicated by the
         data starting from the syntax element general_profile_space to
         the syntax element general_level_idc, inclusive, in the first
         profile_tier_level( ) syntax structure in vpsA can decode any
         bitstream that conforms to the profile, tier, level, and
         constraints indicated by the data starting from the syntax
         element general_profile_space to the syntax element
         general_level_idc, inclusive, in the first profile_tier_level(
         ) syntax structure in vpsB.


  *   I don’t understand what makes the first PTL structure in a VPS special. Could you explain? Perhaps adding an informative NOTE in the text would help readers to understand this design choice.

Point 2 above should take care of this

[YS] This is a carry over from the HEVC payload format which only applied to the single layer-case and therefore the first PTL structure was mentioned there. We have more than on PTL in layered bitstream for VPSs but probably we need to do as suggested below and be stricter on all PTLs in a VPS.

[MH] Right, Yago, the already concluded point 9 should take care of this without treating the first PTL in a VPS in a special way.


8) Section 7.1, related to sprop-sei


  *   It would be helpful to enable offer-answer negotiation on properties that are signaled with SEI messages. These properties could include e.g.:

     *   Whether film grain synthesis is supported (the film grain characteristics SEI message)
     *   Whether frame packing is supported and which frame packing arrangement type to use (the frame packing arrangement SEI message)
     *   How many subpictures to use and which levels they should correspond (subpicture level information SEI message)


  *   It is suggested to allow empty or truncated SEI message payloads in sprop-sei to indicate the capability of the offerer to encode these SEI messages with any SEI message payload that starts with the bits included in sprop-sei (if any). For example, when contained in sprop-sei:

     *   The film grain characteristics SEI message with a truncated payload that only contains data up to fg_model_id indicates that the offerer is capable of producing a bitstream containing film grain of that fg_model_id value.
     *   The frame packing arrangement SEI message with an empty payload indicates that the offerer is capable of producing a bitstream with any frame packing arrangement.
     *   The subpicture level information SEI message with an empty payload indicates that the offerer is capable of encoding content with subpictures and include subpicture level information SEI message(s) into the bitstream.


  *   It is suggested to introduce an optional recv-sei parameter along the following principles:

     *   The SEI messages included in recv-sei in an answer indicate which SEI messages must be present in the bitstream from the offerer to the answerer
     *   When recv-sei is present in an answer, the SEI message types in recv-sei must be the same as or a subset of those included in the sprop-sei of the offer.
     *   When recv-sei is present in an answer, the SEI message payload for a particular SEI message type in recv-sei must start with the SEI message payload bits present (if any) in recv-sei for the same SEI message type.
     *   It is suggested to allow empty or truncated SEI message payloads in recv-sei to indicate that the answerer has the liberty to encode any values for remaining of the SEI message payload (as long as they conform to the specification where the SEI message is specified). Note that the conditions in the bullets above still apply.
     *   When an SEI message with a particular SEI message type is present in sprop-sei of an offer and is absent in recv-sei in an answer, the answerer indicates that it does not support the processing the SEI message included in sprop-sei.

I see the motivation, but implementing any of these above points would be major surgery at the current stage.  Correct me if I’m wrong, but all this could also be “bolted on” later, with no detrimental effect to interop, AFAICT.  I’m a bit reluctant.  Others?


[YS] I agree with Stephan. I understand the motivation but this is quite a big change. In particular, I was wondering whether allowing empty or truncated SEI messages is necessary or whether the SEI manifest SEI message or SEI prefix indication SEI message could be used for this purpose.

[MH] Good point, Yago, on the SEI manifest and SEI prefix indication SEI messages. I think they can indeed be used to avoid the empty or truncated SEI messages in sprop-sei and recv-sei and simplify the design. I’d think no change in sprop-sei is needed and recv-sei is simplified to the following:

     *   The SEI messages included in recv-sei in an answer indicate which SEI messages must be present in the bitstream from the offerer to the answerer
     *   When recv-sei is present in an answer, the SEI message in recv-sei must be one of the following:

        *   the SEI message must be the same as directly included in the sprop-sei of the offer;
        *   the SEI message must have a type indicated in a SEI manifest SEI message in the sprop-sei of the offer;
        *   the SEI message must have a type and start with the respective content indicated by a SEI prefix indication SEI message contained within the srop-sei.

     *   When an SEI message with a particular SEI message type is present in sprop-sei of an offer and is absent in recv-sei in an answer, the answerer indicates that it does not support the processing the SEI message included in sprop-sei.

I continue to think this would be useful for many purposes. I am opening up two of the cases I briefly mentioned above as examples:

  *   The offerer encodes stereoscopic video, with spatial or temporal frame packing. It includes SEI manifest SEI message with frame packing arrangement SEI message in sprop-sei. The answerer can deal with spatial frame packing, but temporal frame packing is not supported by its displaying process. It includes the frame packing arrangement SEI message with a spatial packing type in recv-sei of the answer.
  *   The offerer is able to encode 8K video and supports encoding of independent subpictures. It includes SEI manifest SEI message with subpicture level information SEI message in sprop-sei. The answerer can decoded 8K video only if it uses four 4K-capable processing cores, one processing core per each subpicture sequence. Thus, it includes subpicture level information SEI message indicating four subpictures, each with a level of 4K decoding capacity in the recv-sei of the answer.

10) Section 7.2.2.2, on sprop-sub-layer-id indicating the property of the highest layer only

   o  The offerer MAY include sprop-sub-layer-id which, in case of
      scalable VVC, is interpreted as the highest sub-layer of the
      highest enhancement layer in the OLS indicated by sprop-ols-id.
      The answerer MAY include recv-sub-layer-id which can be used to
      downgrade the sublayer of the highest enhancement layer.  This
      specification does not support downgrading the sub-layer of any
      layers in the OLS that are not the highest layer.
         Informative note: in other words, using this mechanism, an
         answerer can downgrade only the frame rate for the highest
         spatial/quality layer (typically corresponding to the highest
         resolution or bitrate, hence the most complex to decode), but
         not for lower spatial/quality layers.  The answerer must
         support all sublayers for lower layers in the OLS, or reject
         the offer.  That's not a big burden, as the receiver/decoder
         has the option to discard any sublayers it cannot decode,
         irrespective of what is being signalled through offer/answer.


  *   This design choice is not obvious. Note that pruning temporal sublayers from the highest layer but not from the reference layer(s) could have undesirable impacts, such as:

     *   The decoder outputs reference-layer pictures with TemporalId greater than sprop-sub-layer-id. In SNR/spatial scalability these pictures have lower quality/resolution than the pictures of the highest layer and hence cause temporally fluctuating quality/resolution.
     *   If sprop-sub-layer-id is used to input the Htid variable to the decoder as specified in Section 8.1.1 of VVC, the receiver would need to apply the sub-bitstream extraction process of VVC, C.6 to obtain a conforming bitstream.



  *   I would think that the maximum of sprop-sub-layer-id (when present) and recv-sub-layer-id (when present) MUST indicate the highest temporal sublayer to be decoded for the bitstream from the offerer to the answerer and MAY be provided as the value of Htid to the VVC decoding process as specified in Section 8.1.1 of VVC. Changes would be needed in Section 7.2.2.2 as well as in sprop-sub-layer-id and recv-sub-layer-id in Section 7.1.

Umm.  I’m lost.  Yago/Ye-Kui?

[YS] We had some internal discussion before the last update and the point that was made is that it seems that some implementations using scalability use pruning of the higher layers mainly. Having said that I agree with the issues pointed out.

So
1) Yes you would have temporally fluctuating quality/resolution
2)  sprop-sub-layer-id cannot be used to set Htid and the extraction process drops some sublayers of the highest layer, which is not how it is defined in VVC. Basically, sprop-sub-layer-id would not change the operating point and picture from non-output layers need to be output.


So if we keep this design choice we need to clarify this; that we are doing something different than how VVC defines operating points. My recollection is that this is how some products are implemented and differs from how VVC defines operating points. Stephan could you confirm that my recollection is correct?

So not sure how to proceed here. If this is how it is typically implemented, it seems reasonable to do things a bit different than how operating points are defined in VVC, but we should clearly state it. Any views on how to proceed here?

[MH] I am having hard time to understand the motivation why only the highest layer would undergo temporal sublayer pruning. And if so, would the temporally fluctuating quality/resolution be really a desirable way to display such video. Unless there are good answers, I think the I-D should be aligned with the way VVC defines operating points.


Best Regards,
Miska