| Abstract | The rapid growth of multimedia products and services creates strong competition in terms of content quality and user acceptance. While objective quality assessment methods provide useful technical insights, subjective evaluation remains essential for capturing true human perception. Most subjective quality studies are conducted in controlled laboratory environments with a number of viewers adequate to enable statistically significant results. This presents two main limitations: it is difficult and time expensive to involve a large number of participants in subjective tests; not all subjects have the same accuracy in assessing the quality of the content. To address these limitations, an alternative approach is to collect ratings from a smaller group of subjects and apply statistical models to estimate the mean opinion score (MOS) as provided by a large group of viewers. In this paper, several statistical models for MOS estimation are evaluated using small subsets of a subjective dataset. Specifically, EPMOS, based on ordinal logistic regression, is examined alongside the standardized ITU-T P.910, ITU-T P.913, and ITU-R BT.500 models. The estimated results from each model are compared with the reference MOS computed from the full dataset. The results demonstrate that EPMOS provides estimation errors and correlations with the reference scores that closely match the best performance across models, even if it does not always achieve the absolute optimum. The complete set of results is made publicly available upon publication of this paper. |
|---|