Re: [Cellar] AV1 seeking

Andreas Rheinhardt <andreas.rheinhardt@googlemail.com> Sun, 15 July 2018 22:20 UTC

Return-Path: <andreas.rheinhardt@googlemail.com>
X-Original-To: cellar@ietfa.amsl.com
Delivered-To: cellar@ietfa.amsl.com
Received: from localhost (localhost [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id DFDFF130E8C for <cellar@ietfa.amsl.com>; Sun, 15 Jul 2018 15:20:48 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -2
X-Spam-Level:
X-Spam-Status: No, score=-2 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, FREEMAIL_FROM=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_PASS=-0.001] autolearn=ham autolearn_force=no
Authentication-Results: ietfa.amsl.com (amavisd-new); dkim=pass (2048-bit key) header.d=googlemail.com
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id Ey2DgI_sINUT for <cellar@ietfa.amsl.com>; Sun, 15 Jul 2018 15:20:46 -0700 (PDT)
Received: from mail-wr1-x442.google.com (mail-wr1-x442.google.com [IPv6:2a00:1450:4864:20::442]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by ietfa.amsl.com (Postfix) with ESMTPS id 4D894130E88 for <cellar@ietf.org>; Sun, 15 Jul 2018 15:20:46 -0700 (PDT)
Received: by mail-wr1-x442.google.com with SMTP id a3-v6so20795416wrt.2 for <cellar@ietf.org>; Sun, 15 Jul 2018 15:20:46 -0700 (PDT)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=googlemail.com; s=20161025; h=subject:to:references:from:message-id:date:mime-version:in-reply-to :content-transfer-encoding; bh=hRkfIk2PNO8D4kru70fJ6RRUcNWIellb6URYh1ukfKo=; b=AIGic6c3Nqx0GE2ai3uzv7F3tXUIPwjanjVXFOKbykvS10fNbNwSQmNhgzE0pcEDq7 lmKnGBZ2QyHaOzRgRke3wIzW1+bW9ZS71WT8cCeZaTkQO/umYOCL+6TEOfBfItsJ9K+t kDBvHxPxdCLVNVbOJJSrkrdYYdO8g7lUlVGB7LhSKrYjibU8Mj14BF8thQMA2umVgwlp ee6qJSmy0GfUtac+gSEHy+2YoR/TAJWm2NnN6BV1pDUaGxRZSjjczdpdevnzG+7ym48g cB6ogppDRZFhb0IDNGXgv2Gv2dZe2vBim52Na1vMKp/G7naI1ELF3XaodzJlKrj13hLN IKUQ==
X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20161025; h=x-gm-message-state:subject:to:references:from:message-id:date :mime-version:in-reply-to:content-transfer-encoding; bh=hRkfIk2PNO8D4kru70fJ6RRUcNWIellb6URYh1ukfKo=; b=B36X9AkeTjFueVASHqHv/TCzCmi1OvXUU1akDUa23hO1v9OK/qzrmEm/DWV6Y2otJR sthLjZ4aT/elXVGviszITOfhveuua0kWxu7Q4+uSr//jilHh9u2ZEgMdEY2XdPvWtlcS tzihIaY/fYjeYiX8x4lwAxJisEFW1ud44QpgQUoMamCBocfd2QgnIA+pGeCiWk9J5PU0 omq8+SIYBPDLTm9wtlyviKu6YzyxbDRBgLV+jAk/46FJSLVLAmgMRekbd3M83wNEuBvQ BGbNVV2SmQDgBTP+DfuR9aOA41SN22OIWf3mxSnnHgJhPfsXYtTFCE1V0/iLYVUN6f2B eHag==
X-Gm-Message-State: AOUpUlGEe5JL3GnPavkwhVoTWf9Y/YNYP0oRuCJ/dbstS2ouP6faaryM PpFcOWL0qLhc4cCZy1H2Go4k25Me
X-Google-Smtp-Source: AAOMgpenuIVuKCSU1gyrvr4jbrqhmDEQORGUqETI0OKcgedA53IHYxhYCoq7xJMRLT10luJyAv1nzw==
X-Received: by 2002:adf:e491:: with SMTP id i17-v6mr10868606wrm.145.1531693244508; Sun, 15 Jul 2018 15:20:44 -0700 (PDT)
Received: from [127.0.0.1] ([2a00:1298:8011:212::165]) by smtp.googlemail.com with ESMTPSA id l7-v6sm10759186wmh.1.2018.07.15.15.20.43 for <cellar@ietf.org> (version=TLS1_2 cipher=ECDHE-RSA-AES128-GCM-SHA256 bits=128/128); Sun, 15 Jul 2018 15:20:43 -0700 (PDT)
To: cellar@ietf.org
References: <CAOXsMFKTNCxYcviYS0h_VYjegV3RZFvZ7AV7GhdCq=oeGmgMuQ@mail.gmail.com>
From: Andreas Rheinhardt <andreas.rheinhardt@googlemail.com>
Message-ID: <62c29889-49f5-6634-049a-a2d73315bb3c@googlemail.com>
Date: Sun, 15 Jul 2018 22:19:00 +0000
MIME-Version: 1.0
In-Reply-To: <CAOXsMFKTNCxYcviYS0h_VYjegV3RZFvZ7AV7GhdCq=oeGmgMuQ@mail.gmail.com>
Content-Type: text/plain; charset="utf-8"
Content-Transfer-Encoding: 7bit
Archived-At: <https://mailarchive.ietf.org/arch/msg/cellar/2ZEHIxpFcLZBIlpHn0JOqOKBtHA>
Subject: Re: [Cellar] AV1 seeking
X-BeenThere: cellar@ietf.org
X-Mailman-Version: 2.1.27
Precedence: list
List-Id: Codec Encoding for LossLess Archiving and Realtime transmission <cellar.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/cellar>, <mailto:cellar-request@ietf.org?subject=unsubscribe>
List-Archive: <https://mailarchive.ietf.org/arch/browse/cellar/>
List-Post: <mailto:cellar@ietf.org>
List-Help: <mailto:cellar-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/cellar>, <mailto:cellar-request@ietf.org?subject=subscribe>
X-List-Received-Date: Sun, 15 Jul 2018 22:20:49 -0000

Hello,

Steve Lhomme:
> Using BlockReference we can actually know where to seek for these
> particular frames to get all the frames they need. But most people use
> SimpleBlock (or we could forbid its use for such streams) and it would
> mean referencing all frames in the Cues which is not a good idea.
> 
I actually have already made some proposals that allow the demuxer to
know where to seek to. Here are two old ones and a new one (c)) which is
my new favourite:

a) One could reference the recovery points and the keyframe (i.e.
non-delayed) RAPs in the cues. Both would be normal cue entries. The
recovery points would have to be put into a `Block` in a `Blockgroup`
and said `Blockgroup` must have a `ReferenceBlock` that points to the
delayed RAP.
Pro: Normally the delayed RAP and the recovery point end up in the same
cluster and so hopefully the delayed RAP can be quickly accessed because
it is still cached. Also this approach doesn't rely on anything
deprecated in Matroska or unavailable in Webm. It can be generalized to
the other gradual decoder refresh scenarios.* It has very low overhead.
Con: Depending on the demuxer implementation this might result in two
seeks; it is also not really very backward-friendly: If one has a
delayed RAP at t_0, the recovery point at t_1 and the next RAP at t_2
with t_0<t_1<t_2 and a user wants to seek to somewhere between t_1 and
t_2, a player that doesn't support this kind of seeking will seek to t_1
and either refuse to decode said frame, because it isn't a keyframe, (in
which case it probably proceeds to t_2 and starts decoding there) or
will decode resulting in corrupted output.

b) One could reference both recovery points and keyframes. The
`CuePoint` (actually CueTrackPositions) for recovery points would
contain `CueReference` containing `CueRefTime` (whose value would be the
timestamp of the (Simple)Block containing the delayed RAP) and `CueCluster`.
Pro: Only one seek is necessary. Can be generalized (actually it is
already general) to the general case of gradual decoder refresh. Low
overhead.
Con: Uses elements that are deprecated in Matroska and unavailable in
Webm. A demuxer that doesn't handle these uncommon elements will have
the same problem as in a).

c) To seek adaequately the demuxer doesn't need to know the complete
reference structure. Actually, experience has shown that this
information is both unneeded and undesired at the container level. The
only things that are needed are the timestamp and position of the RAP
and the timestamp of the block from which point on the output is
considered acceptable.
Therefore I propose to create a new element (child of
`CueTrackPositions`) called `DecoderConvergence`/`OutputGood` or
whatever (I don't have a good name for it -- suggestions welcome!); this
element (a signed (variable sized) integer) would contain the timestamp
offset (in TimestampScale units) relative to the CuePoint`s
corresponding `CueTime` from which on the output of the decoder is
considered good/acceptable when one starts decoding at the referenced
block. No default value and if this element is not present then nothing
can be inferred regarding the point from which on the output will be
acceptable. (We shouldn't simply make the element mandatory with default
value 0, because there are already intra-decoder-refresh H.264 files and
muxers (like mkvmerge) have simply referenced the frames that contain a
recovery point SEI even though outputting said frame is unacceptable
because they are damaged (when decoding began at said frame).)
The AV1 usage is clear: The delayed RAPs that are later output at
recovery points get referenced and the offset of the accompanying
recovery point ends up in the new element. Similar for
intra-decoder-refresh scenarios.
There is also another usage: It can be used for H.265's decodable
leading (RADL) pictures: This is like open-gop, but with the difference
that the frames that follow the keyframe in decoding order and precede
it in display order don't reference anything that precedes the keyframe
in decoding order so that these frames can be correctly decoded. This
means that you can effectively use the "golden" keyframe as a suitable
reference for more frames. This is why the element is signed.
(This also shows that `DecoderConvergence` is not a fitting name. As has
been said: Suggestions welcome!)

Pro: Low overhead. Backward compatible with demuxers that don't
understand the element (as long as they ignore it as they should).
General solution. Only one seek is necessary.
Con: Requires a new element, in particular we need to talk to the Webm
guys; doesn't solve the problem for files without cues.


*: In this case the RAP where decoding starts might not be one of the
frames that the frame where the output is good directly references; yet
it indirectly references it, because the direct references reference it
(or the references of the references; you get the idea). Nevertheless
one should add a `ReferenceBlock` pointing to said RAP (and this is the
only needed `ReferenceBlock`). But I have to add that this is a
deviation from the current understanding of the `ReferenceBlock` element.

> In MP4 they have [initial_presentation_delay_minus_one] in the
> CodecPrivate. I did not understand it so far because it's not found in
> the AV1 spec. But it seems to guarantee that to read frame 'f' you
> need to decode X frame before that one. In our case that would be 5 to
> have at least a decoded.
> 
You completely misunderstood this field. It is the equivalent of the
max_num_reorder_frames value from H.264. Let me explain it in MPEG
terminology (with b-frames) as you probably have way more experience
with this: Consider a decoder that can only decode one frame per unit of
time and a stream like this (left to right is decoding order; the
numbers are presentation order):
I0 P2 B1 ...
If one displayed the leading I frame immediately after decoding it, one
would not have the right frame to display at time 1, because at that
time only I0 and P2 has been decoded, not B1. Therefore one has to
decode I0 and P2 before one outputs the first frame and and
max_num_reorder_frames would be 1. The typical b-pyramid would require
to decode the first three frames before the display of the first frame
and max_num_reorder_frames would be 2.
The [initial_presentation_delay_minus_one] is the AV1 analogue of this.
This number is not an upper bound for the amount of temporal units
between delayed RAP and recovery point/for the amount of frames shared
between a GOP. Just look at this example:
MPEG example
I0 P5 B1 B2 B3 B4 I10 B6 B7 B8 B9
AV1 example (I kept the MPEG-naming with P and B to make it easier
comparable to the above; furthermore the pointer *x denotes that a frame
is showable and x without * is a frame header that outputs *x via
show_existing_frame;  square brackets are the delimiters of temporal
units; I10 is the delayed RAP frame)
[I0] [*P5 B1] [B2] [B3] [B4] [P5] [*I10 B6] [B7] [B8] [B9] [I10] ...
This stream can have [initial_presentation_delay_minus_one] equal to 1,
yet in order to seek to [I10] one has to decode the temporal unit [*I10
B6] (or at least the decodable keyframe in it) which is four temporal
units in front of [I10].

> IMO this is a bad design because it mixes the container with the codec
> (something done in the past in ogg with bad results). Seeking should
> not ask the codec how it should be done. Also that means to mux such
> an MP4 you'd need to scan the whole source first to verify the value
> is correct or that information needs to be given to the muxer by the
> encoder/packetizer (this information is not found in the stream).
> 
> We have the option of not caring as it's allowed to leave it in grey
> area by the AV1 spec. Or we can think of a proper "container"
> solution. ie storing this information in the TrackInfo but not in the
> CodecPrivate. We already have a CodecDelay but IMO it doesn't fit the
> bill. Each frame has the timestamp modified by it. Which is not the
> case here. We may need a SeekingDelay so that looking for a seek point
> takes this amount in account.
> 
yes, this is completely different than CodecDelay. There is a Matroska
field that could be used for this: SeekPreRoll. But there are two
problems with this: This value is valid for the whole track, whereas the
amount of time between delayed RAP and recovery point can vary. And of
course the needed value is generally unknown at the beginning of the
muxing process. Not to mention the fact that actually SeekPreRoll is
valid if one starts decoding at an arbitrary frame, not necessarily a
(delayed) keyframe. Therefore a valid SeekPreRoll element for the track
would have to be so large that it would be impractical.

- Andreas Rheinhardt