[ippm] Re: Responsiveness under working conditions

Joachim Fabini <Joachim.Fabini@tuwien.ac.at> Mon, 02 December 2024 09:20 UTC

Return-Path: <joachim.fabini@tuwien.ac.at>
X-Original-To: ippm@ietfa.amsl.com
Delivered-To: ippm@ietfa.amsl.com
Received: from localhost (localhost [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id D9C8AC151093 for <ippm@ietfa.amsl.com>; Mon, 2 Dec 2024 01:20:40 -0800 (PST)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: 3.093
X-Spam-Level: ***
X-Spam-Status: No, score=3.093 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, GB_SUMOF=5, RCVD_IN_DNSWL_BLOCKED=0.001, RCVD_IN_ZEN_BLOCKED_OPENDNS=0.001, SPF_PASS=-0.001, T_SCC_BODY_TEXT_LINE=-0.01, URIBL_DBL_BLOCKED_OPENDNS=0.001, URIBL_ZEN_BLOCKED_OPENDNS=0.001] autolearn=no autolearn_force=no
Received: from mail.ietf.org ([50.223.129.194]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id J7UsgeIcgePL for <ippm@ietfa.amsl.com>; Mon, 2 Dec 2024 01:20:37 -0800 (PST)
Received: from secgw2.intern.tuwien.ac.at (secgw2.intern.tuwien.ac.at [IPv6:2001:629:1005:30::72]) (using TLSv1.2 with cipher ECDHE-ECDSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by ietfa.amsl.com (Postfix) with ESMTPS id EEE30C14F5EA for <ippm@ietf.org>; Mon, 2 Dec 2024 01:20:35 -0800 (PST)
Received: from Kiteworks (kwmta2.intern.tuwien.ac.at [128.130.30.92]) by secgw2.intern.tuwien.ac.at (8.14.7/8.14.7) with ESMTP id 4B29KVLw028384; Mon, 2 Dec 2024 10:20:31 +0100
Received: from secgw2.intern.tuwien.ac.at ([128.130.30.72]) by totemomail.intern.tuwien.ac.at (Totemo SMTP Server) with SMTP ID 651; Mon, 2 Dec 2024 09:20:31 +0000 (GMT)
Received: from edge19b.intern.tuwien.ac.at (edge19b.intern.tuwien.ac.at [IPv6:2001:629:1005:30::46]) by secgw2.intern.tuwien.ac.at (8.14.7/8.14.7) with ESMTP id 4B29KVfJ028377 (version=TLSv1/SSLv3 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=FAIL); Mon, 2 Dec 2024 10:20:31 +0100
Received: from mbx19c.intern.tuwien.ac.at (2001:629:1005:30::83) by edge19b.intern.tuwien.ac.at (2001:629:1005:30::46) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.2.1544.13; Mon, 2 Dec 2024 10:20:30 +0100
Received: from [IPV6:2001:871:222:b8ad:7df9:2b2a:748b:9e4e] (2001:871:222:b8ad:7df9:2b2a:748b:9e4e) by mbx19c.intern.tuwien.ac.at (2001:629:1005:30::83) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.2.1544.13; Mon, 2 Dec 2024 10:20:30 +0100
Message-ID: <6746096c-c5d3-4021-b53b-6552b0c71171@tuwien.ac.at>
Date: Mon, 02 Dec 2024 10:20:29 +0100
MIME-Version: 1.0
User-Agent: Mozilla Thunderbird
To: Ruediger.Geib@telekom.de, chris.box.ietf@gmail.com
References: <CACJ6M16-2y9RT15SooPVZjfuVY+2-LOP3menYf6DR6reaJh1SQ@mail.gmail.com> <BEZP281MB200764B98BE27D6BBE55176D9C582@BEZP281MB2007.DEUP281.PROD.OUTLOOK.COM>
Content-Language: en-US
From: Joachim Fabini <Joachim.Fabini@tuwien.ac.at>
In-Reply-To: <BEZP281MB200764B98BE27D6BBE55176D9C582@BEZP281MB2007.DEUP281.PROD.OUTLOOK.COM>
Content-Type: multipart/signed; micalg="pgp-sha256"; protocol="application/pgp-signature"; boundary="------------35FG0wwvH0E9uIoj8B11i1Ly"
X-ClientProxiedBy: mbx19d.intern.tuwien.ac.at (2001:629:1005:30::84) To mbx19c.intern.tuwien.ac.at (2001:629:1005:30::83)
Message-ID-Hash: TTXAVYEJO4N4LQRAIWKBBXKBW6NHTXHM
X-Message-ID-Hash: TTXAVYEJO4N4LQRAIWKBBXKBW6NHTXHM
X-MailFrom: joachim.fabini@tuwien.ac.at
X-Mailman-Rule-Misses: dmarc-mitigation; no-senders; approved; emergency; loop; banned-address; member-moderation; header-match-ippm.ietf.org-0; nonmember-moderation; administrivia; implicit-dest; max-recipients; max-size; news-moderation; no-subject; digests; suspicious-header
CC: ippm@ietf.org
X-Mailman-Version: 3.3.9rc6
Precedence: list
Subject: [ippm] Re: Responsiveness under working conditions
List-Id: IETF IP Performance Metrics Working Group <ippm.ietf.org>
Archived-At: <https://mailarchive.ietf.org/arch/msg/ippm/9JuY0xzuv9l1JSC_z9BvT-KstIA>
List-Archive: <https://mailarchive.ietf.org/arch/browse/ippm>
List-Help: <mailto:ippm-request@ietf.org?subject=help>
List-Owner: <mailto:ippm-owner@ietf.org>
List-Post: <mailto:ippm@ietf.org>
List-Subscribe: <mailto:ippm-join@ietf.org>
List-Unsubscribe: <mailto:ippm-leave@ietf.org>

Hi Chris, Hi Rüdiger

I'd add some (technical and partly philosophical) aspects to the 
observations of Rüdiger. Some topics worth considering complicate 
matters even more:

1. Traffic & loss patterns: Metrics like loss period and loss distance 
(RFC3357) have a huge impact on the perceived QoS. One can easily design 
streams having identical loss percentage but fundamentally distinct QoE. 
Isolating QoE observations from application-layer information (like the 
codec - how good it is in concealing the losses) is nearly impossible. 
Not even mentioning the difference in terms of QoE when losing i-frames 
vs. losing p-frames within a video stream. So let's stick to 
network-layer measurements.

2. On-demand capacity allocation is key to the scalability of today's 
mobile cellular access networks. Additional uncertainty factors like 
users-in-a-cell, capacity-shared-within-a-cell, etc. contribute to the 
network state and render measurements in a mobile network a one-time 
experience. We can attempt to restrict uncertainty factors - but this 
limits the representativity of measurements to the specific scenario.

3. More then ten years ago, Al Morton and I have summarized some of 
these observed (and confirmed) peculiarities in RFC 7312 - in particular 
considerations on *repeatability*, *continuity*, and *conservative*. 
These three are closely related to the terms that you used, 
repeatability and relevance. You can find more details by searching for 
the contributions to IRTF RAIM 2015 (slides including measurement 
results and references to publications at various conferences and journals).

Key to the RFC2330-defined term repeatability is the concept of 
"identical conditions". Today's real networks are stateful and much too 
complex (i.e. involve too many uncertainty factors that 
users/measurements can not control). Main metric for vendors and 
operators is overall network capacity optimization (=revenue) - at the 
cost of high non-determinism from a single user's perspective. I'm 
convinced that now and in the future we can and will *never* safeguard 
identical conditions even for network-layer measurements (except for 
fully isolated measurements, or unless the principles of network design 
change fundamentally). With application-layer semantics assigned to 
specific packets (some are more "valuable" than others) this becomes 
even more complex.

So I'd like to stress that repeatability is a purely theoretical concept 
that measurements should aim at. I do not say that we should refrain 
from designing better measurement methods, on the contrary. But my 
impression is that we're attempting/expecting the impossible. The 
design, implementation and configuration of today's network technologies 
is the reason why neither users nor measurements can anticipate or 
extrapolate network behavior. What measurements can do is to assess how 
the network behavior was for one specific stream or traffic under 
consideration that was monitored. Any inference beyond this is 
fortunetelling (imho slightly beyond the scope of the IPPM charter): 
this includes in particular inference on the behavior/results of the 
same network at a distinct point in time, on the experience of other 
users, or on the network behavior of a distinct stream of the same user.

One valid conclusion is, however, to request network vendors and 
operators to improve the determinism of their networks. Ultimately this 
shortcoming is the main challenge, both to users and measurements.

kind regards,
Joachim



On 11.11.24 17:10, Ruediger.Geib@telekom.de wrote:
> Hi Chris,
> 
> capturing QoE in a simple and comprehensive way which is linked to a 
> what a person perceives just experiencing performance issues isn’t 
> simple, I think.
> 
> Sure, offering two or three dimensions minimum delay, buffer caused 
> latency and packet loss helps. Looking at the work invested by ITU-T in 
> MOS modeling, I’d by surprised if any simple metric can be built. Voice 
> suffers more from serious jitter than from low percentage drop. Other 
> apps may suffer more from packet drop than they do from jitter. During 
> the IETF meeting, I noted that slides and talk for remote presentations 
> are more important, than slides and presenter picture. I think, models 
> for interactive remote control start to be built and there picture 
> quality degradation may be more acceptable, than freeze.
> 
> I appreciate efforts to improve situation, but as I mentioned, there may 
> be no simple one-approach-fits-all solution.
> 
> The ITU-T recommendation mentioned by you is likely Y.1541, 
> https://www.itu.int/itu-t/recommendations/rec.aspx?rec=11462 
> <https://www.itu.int/itu-t/recommendations/rec.aspx?rec=11462>. I’d 
> guess later ITU-T work on streaming and voice QoE is more relevant. 
> Examples are
> 
> P.863: Perceptual objective listening quality prediction, I think, and
> 
> P.1200-P.1299: Models and tools for quality assessment of streamed media.
> 
> Work related to VR and modeling of interactive services is ongoing in ITU-T.
> 
> Regards,
> 
> Rüdiger
> 
> *Von:* Chris Box <chris.box.ietf@gmail.com>
> *Gesendet:* Samstag, 9. November 2024 10:25
> *An:* ippm@ietf.org
> *Betreff:* [ippm] Responsiveness under working conditions
> 
> Hi everyone
> 
> I watched Stuart's presentation on Monday and I'm personally convinced 
> that getting this responsiveness test right is one of the most important 
> tasks of internet engineering right now. Responsiveness is the next 
> frontier of quality improvement, and the current experience of most is 
> plentiful bandwidth but frequently poor responsiveness. We can do 
> better, and this test is the key enabler for that.
> 
> My personal aims sound very similar to Stuart's:
> 
>     Primary goals (must have): Relevance and Repeatability
> 
>     Secondary goal (desirable): Convenience
> 
> When I heard Jonathan describe the once-per-minute wifi freeze, I 
> concluded this is a case where we ought to sacrifice convenience in 
> favour of relevance and repeatability. If end users are being impacted 
> by that 500ms gap, then the test ought to measure that. It's much less 
> helpful if it ignores that degradation. Of course periodic (or even 
> sporadic) disruptions can occur on longer timescales and we have to draw 
> a line somewhere. But perhaps we can consider a "full test" and a "quick 
> test" mode.
> 
> The other major open question is how to convert from a large set of 
> samples to a single RPM value. I agree with Abhishek that the message 
> from Bjorn's QoO distribution is that the outliers have the most 
> significant effect on user experience. Those little black circles are 
> the ones we notice the most. So I suggest we do not discard any values. 
> They all count. My view is that the packet with the maximum delay should 
> form a significant component of the RPM. It's not everything of course, 
> and repeatability probably implies we need some way to compute a "worst 
> delay" figure from multiple packets, e.g. the worst 10 or 20.
> 
> ITU Recommendation Y.1514 defines network performance requirements. It 
> says for class 0 audio over a measurement period of one minute, the mean 
> IP Packet Transfer Delay should not exceed 100ms, and IP Packet Delay 
> Variation (RFC3393) should not exceed 50ms. To have good enough 
> responsiveness for a video conference, this means mean delay in each 
> direction <100ms, and no audio-bearing packets should arrive later than 
> 150ms. With responsiveness we're measuring the sum of both directions, 
> so it will depend on the assymetry of delay, for example in a mobile 
> network uplink packets need to wait until granted permission. What we 
> can say is that for all networks, http_l of <100 mean and <150 max is 
> sufficient to meet the class 0 requirement. For a symmetrical network, 
> http_l of <200 mean and <300 max is sufficient.
> 
> Also ITU G.114 discusses delay due to the use of an IP delay variation 
> buffer. It says that for planning purposes, it is recommended to assume 
> that a de-jitter buffer adds one half of its peak delay to the mean 
> network delay.  A jitter buffer designed to compensate for 50 ms packet 
> delay variation range will introduce 25 ms additional delay, on average.
> 
> So what should we compute the RPM from? I suggest it needs to be some 
> combination of the worst delay and the typical delay, that in testing 
> proves itself to be both Relevant and Repeatable.
> 
> Chris
> 
> 
> _______________________________________________
> ippm mailing list -- ippm@ietf.org
> To unsubscribe send an email to ippm-leave@ietf.org