Deep CNN Model for Non-Screen Content and Screen Content Image Quality Assessment

Send Message

To: Author

Deep CNN Model for Non-Screen Content and Screen Content Image Quality Assessment

Article Fingerprint

ReserarchID

3I95A

Deep CNN Model for Non-Screen Content and Screen Content Image Quality Assessment Banner

Key Research Insights

Synthesized scholarly intelligence & interactive research assistant
  • English
  • Afrikaans
  • Albanian
  • Amharic
  • Arabic
  • Armenian
  • Azerbaijani
  • Basque
  • Belarusian
  • Bengali
  • Bosnian
  • Bulgarian
  • Catalan
  • Cebuano
  • Chichewa
  • Chinese (Simplified)
  • Chinese (Traditional)
  • Corsican
  • Croatian
  • Czech
  • Danish
  • Dutch
  • Esperanto
  • Estonian
  • Filipino
  • Finnish
  • French
  • Frisian
  • Galician
  • Georgian
  • German
  • Greek
  • Gujarati
  • Haitian Creole
  • Hausa
  • Hawaiian
  • Hebrew
  • Hindi
  • Hmong
  • Hungarian
  • Icelandic
  • Igbo
  • Indonesian
  • Irish
  • Italian
  • Japanese
  • Javanese
  • Kannada
  • Kazakh
  • Khmer
  • Korean
  • Kurdish (Kurmanji)
  • Kyrgyz
  • Lao
  • Latin
  • Latvian
  • Lithuanian
  • Luxembourgish
  • Macedonian
  • Malagasy
  • Malay
  • Malayalam
  • Maltese
  • Maori
  • Marathi
  • Mongolian
  • Myanmar (Burmese)
  • Nepali
  • Norwegian
  • Pashto
  • Persian
  • Polish
  • Portuguese
  • Punjabi
  • Romanian
  • Russian
  • Samoan
  • Scots Gaelic
  • Serbian
  • Sesotho
  • Shona
  • Sindhi
  • Sinhala
  • Slovak
  • Slovenian
  • Somali
  • Spanish
  • Sundanese
  • Swahili
  • Swedish
  • Tajik
  • Tamil
  • Telugu
  • Thai
  • Turkish
  • Ukrainian
  • Urdu
  • Uzbek
  • Vietnamese
  • Welsh
  • Xhosa
  • Yiddish
  • Yoruba
  • Zulu
Reading Preferences
Font Size
Line Spacing
Background
This converted HTML version may contain rendering inconsistencies. Please refer to the PDF for the authoritative version, or click here to provide feedback.

Abstract

In the current world, user experience in various platforms matters a lot for different organizations. But providing a better experience can be challenging if the multimedia content on online platforms is having different kinds of distortions which impact the overall experience of the user. There can be various reasons behind distortions such as compression or minimal lighting condition while taking photos. In this work, a deep CNN-based Non-Screen Content and Screen Content NR-IQA framework is proposed which solves this issue in a more effective way. The framework is known as DNSSCIQ. Two different architectures are proposed based upon the input image type whether the input is a screen content or non-screen content image. This work attempts to solve this by evaluating the quality of such images

I. INTRODUCTION

Image quality assessment is a subject of extensive analysis over the last four decades. Different multimedia applications streaming images and videos like Netflix, Amazon Prime Video, Twitter, Face book, Share Chat, etc. are gaining more popularity day by day. With the increasing availability of Internet all over the world, the usage of these applications is increasing rapidly. So, these applications requires quality assessment to be done on their content so that they can provide quality content on their platform. This helps to improve customer visual experience on their respective plat- forms. The main aim of image quality assessment is to quantitatively measure the perceived quality of digital and natural photographs. The acquisition, transmission, storage, post-processing, or compression of images brings different distortions, such as Gaussian blur (GB), Gaussian white noise (WN), or blocking artifacts. WN is added while taking pictures at night with a mobile, GB occurs if not focusing correctly before taking the shot.

Based on IQA results, decisions can be taken on compression ratio for these digital images before storing them in servers for streaming purpose as well as deciding which image will be good to be published on the online platform. A dependable IQA technique can help assess the quality of photos downloaded from the web, as well as measure the accuracy of image processing techniques precisely, such as super-resolution and image compression from a human's perspective. The IQA algorithms are categorized into 3 groups, based upon the usage of reference image: no reference IQA (NR-IQA), reduced-reference IQA (RR-IQA) and full-reference IQA (FR-IQA). The performance of these algorithms is NR-IQA, RR-IQA, and FR-IQA, in order of increasing accuracy. However, since pristine images are not available in most of the real time situation, NR-IQA is most suitable method. The image quality assessed using no-reference (NR) IQA algorithms does not require knowledge of the original image. The image quality assessed using reduced-reference (RR) IQA methods requires only a few details about the original image. Full-reference (FR) algorithms need both a distorted image and a reference image as input and produce a quality rating for the distorted image in comparison to the original image. The most common technique to FR-IQA is to first calculate the local pixel-wise differences between reference image and distorted image. Finally, combine these local calculations into a single scalar value to represent the overall quality difference. Example of FR-IQA algorithms are: Structural Similarity Index Mean (SSIM), the peak signal-to-noise ratio (PSNR) and mean-squared error (MSE). Unlike FR-IQA, in NR-IQA the quality is measured using the features obtained from the distorted images and the subjective quality scores.

This section provides a brief detail of the existing no-reference and reference image quality assessment techniques. Li et al. [1] proposed a new multiscale directional transform, basically a shearlet transform used to extract simple features from distorted images. Then these primary features are used to explain the nature of original images and distorted images.

Then, stacked autoencoders are used to amplify the primary features and make them more distinguishable.

Mittal et al. [2] proposed a NSS-based distortion-generic IQA model. This model works best in the spatial domain. BRISQUE does not calculate the distortion-specific features, such as blur, blocking, or ringing. Rather, it uses scene statistics of locally normalized luminance coefficients to quantify losses of naturalness in the image.

Li et al. [3] trained a general regression neural network (GRNN) to assess the quality of image, relative to the human subjective opinion, across a diverse range of distortion types. The features used for assessing the quality of the image include gradient of the distorted image, entropy of phase congruency image, mean value of the phase congruency image, and entropy of the distorted image.

Moorthy and Bovik [4] introduced DIIVINE (Distortion Identification-based Image Verity and INte grity Evaluation). This algorithm evaluates the quality of a distorted image without the original images. It is a 2-stage based technique where image distortion identification is done first and then image quality assessment is done based on distortion type.

Tang et al. [5] presented a framework, where potentially neither the degradation process nor the ground truth image is known. The method is based on a set of low-level image features. The image quality characteristics are derived from original image measurement and texture statistics. Here, a machine learning technique is used to learn a mapping from these features to the subjective quality scores.

Doermann et al. [6] obtained the basic feature set by the extraction of local features. Then, using the features from the CSIQ database, by adopting K-means clustering, the codebooks with 100 centers was retained. In the mean time, the method proposes high order features: variance, mean, and skewness. The input features are used to get distances to K clusters. Then the method performs regression over three distances. It is sensitive to diverse distortion types.

Fang et al. [7] proposed a quality assessment methodology based on statistical structural and luminance features (NRSL). The evaluations were done on 4 synthetically and 3 naturally distorted image datasets. In terms of high correlation with human subjective judgments, the employed NRSL metric compares favorably to relevant BIQA models. Support vector regression was used to establish the complex nonlinear relationship between feature space and quality score. It was unable to use NRSL for various distortions in chromatic component of the image.

Kim and Lee [8] proposed Deep Image Quality Assessment (DeepQA) where the behavior of HVS is analyzed from the data distribution of IQA datasets. The sensitivity maps were evaluated for various distortion types and degrees of distortion. Subjective score requires reference images.

Y. Li et al. [9] proposed SESANIA where shearlet transform and deep neural networks (stacked autoencoders) is used instead of conventional regression machines. This framework is enhanced to calculate the quality of image in local regions. Liu, Weijer, and Bagdanov [10] used Siamese Network for ranking images in order of image quality. The relative image quality is known for which synthetically generated distortions are used. This helps to solve the issue of the limited size of the IQA dataset. These ranking image sets can be constructed automatically without the requirement of painful effort of labeling by human. This technique uses synthetic images. Saad et al. [11] introduced a Natural Scene Statistics (NSS) based methodology which uses discrete cosine transform (DCT) technique. This method was based on a Bayesian technique to evaluate the image quality scores when features retrieved from the image is given.

Kede Ma et al. [12] proposed an optimized neural network for assessing blind image quality. First, distortion is identified and then the quality prediction is done using the features obtained during distortion identification.

Fei Gao et al. [13] proposed Deep Similarity for image quality assessment (Deep Sim) framework. First, the features of the original and tested images are received from Image Net pretrained VGGNet without any further training. Then, the local similarities between the features of those corresponding images are calculated. At last, the local quality indices are eventually pooled altogether to evaluate the quality index.

Min et al. [14] proposed the concept of multiple pseudo reference images, which are generated from distorted images by applying various levels of distortion. As a result, the quality of a pseudo reference image (PRI) is generally lower than that of its distorted counterpart. The idea behind this methodology is to generate a series of PRI by further degrading the distorted image, and then use local binary patterns (LBP) to calculate the similarity between them to evaluate its quality.

Talebi and Milanfar [15] proposed a convolutional neural network based methodology known as NIMA which is used to predict the distribution of human opinion scores. The network may be used to score images in a way that closely resembles human perception. Its goal is to forecast image technical and aesthetic attributes.

Hou et al. [16] proposed a blind IQA that directly learns qualitative evaluation and predicts scalar values for general usage and fair comparison. Here, the natural scene statistics features are used to represent the images. A discriminative model is trained to distinguish the characteristics into five ranks, that correlate with five rational notion, i.e., bad, poor, fair, good and excellent.

Bose et al. [17] proposed a neural network based method for IQA that enables feature learning and regression in an end-to-end framework. A siamese network using CNN is used with both original and distorted images as input for FR-IQA whereas one branch of siamese network is discarded where the distorted image is used as input for NR-IQA. It incorporates a weighted average patch aggregation that implements a method for pooling local patch qualities to global image quality.

Based on selected feature similarity and ensemble learning, Hammou et al. [18] suggested an ensemble of gradient boosting (EGB) measure. To characterise the perceptual quality distance between the pristine and distorted/processed images, the features obtained from various layers of deep CNN are analyzed. Kang et al. [19] proposed a compact CNN for calculating image quality and identifying distortions. The parameter reduction at the fully connected layers makes this model less prone to overfitting.

III. MOTIVATION

The main motivation behind image quality assessment is to quantify visual perception of humans for image quality so that quality evaluation of images can be done. Digital images intend to degrade during the process from generation to consumption. Different kind of distortions are introduced in the process of transmission, post processing, or compression of images such as white noise, Gaussian blur, or impeding artifacts. This affects the visual experience of users while seeing image content on various online websites. A depend-able IQA algorithm can assist in quantifying the quality of images acquired from the web and also helps to measure the performance of image processing algorithms precisely, such as image-compression and super-resolution, from the point view of a human.

a) Drawbacks of Using CNNs to NR-IQA

Because of its high representation capability and improved performance, convolutional neural networks are the most popular type of neural networks for working with image data. The quantity of the training dataset has a major impact on the performance of neural networks. However, compared to the most frequent computer vision dataset, the currently available IQA datasets are substantially smaller. In contrast to classification datasets, IQA datasets necessitate a time-consuming and sophisticated psychometric experiment. Various data augmentation techniques, such as horizontal reflection, rotation, and cropping, can be employed to enhance the size of the training dataset. The human visual system's (HVS) perception process is made up of several complex processes. It makes training a deep learning model more difficult with a limited dataset. The visual sensitivity of the HVS changes with the spatial frequency of stimuli, and texture prevents concurrent picture alterations.

b) Applications of IQA

IQA has a diverse variety of computer vision and image processing usage. For example:

  • For quantization, an image compression algorithm can use quality as an optimization parameter.

  • Image transmission systems can be created to assess quality and distribute different streaming resources accordingly.

  • Image recommendation algorithms can be created to rank photos according to perceptual image quality.

  • Depending on the image quality desired, several device characteristics for digital cameras can be modified.

IV. PROBLEM STATEMENT

Image Quality Assessment is different from other image processing applications. Unlike segmentation, object detection or classification, preparing IQA dataset is time-consuming and requires complicated psychometric experiments. Therefore, the generation of huge datasets is costly because it requires the supervision of experts which are responsible of ensuring the correct implementation of the experiments. The next drawback is that data augmentation is not preferred because the pixel structure of original images must not be changed. In this paper, an image quality assessment model is developed to calculate the quality of blind images. The distorted images and their ground-truth subjective scores are used for training the CNN model.

V. METHODOLOGY

a) Image Normalization

Image normalization is required because it ensures that the data distribution of each input pixel in the image is consistent. This aids in convergence while doing the training of the neural network. The mean is subtracted from each pixel value, and the result is divided by the standard deviation. Such data would be distributed in a Gaussian distribution centered at zero. The pixel numbers for image input must be positive. As a result, the normalized data must be scaled in the range [0,1] or [0,255]. First, preprocessing is done where the input images are transformed into grayscale, and then they are reduced from their low-pass filtered images. The low-frequency image is retrieved by downscaling the input image to 1/4 and upscaling it again to the original image size. A Gaussian low-pass filter along with subsampling was used to resize the images. The reasons for this kind of normalization is that image distortion doesn't affect the low-frequency component in images. For instance, GB removes high-frequency details, white noise (WN) introduces random high-frequency components to images, and blocking artifacts introduces high-frequency edges. The distortions caused by JPEG is due to excessive image compression. The human visual sensitivity (HVS) is not sensitive to a change in the low-frequency component of the image. The sensitivity reduces rapidly at low frequency.

There is the possibility of losing information while applying a normalization scheme.

b) Architecture

Model for Non-Screen Content IQA: Here a blind image quality assessment method based on CNN is proposed. The features from the CNN are used for a final quality prediction. The design of the network resembles the design of VGG-16 network. The architecture of CNN for synthetic distortion is shown below in figure 1. The existing dataset consists of a subjective score for each distorted image. The model is fine-tuned to evaluate the subjective scores once the training of neural network is completed with enough training data set. The proposed model is fine-tuned on target subject-specific datasets using a variation of stochastic gradient descent.

The kernel size of the convolutions is 3 × 3 . A kernel size of two is used in order to diminish the spatial density in both directions by half. The nonlinear activation function ReLU is used. The feature activations of the final convolution layer's are averaged globally across spatial locations. At the end of the network, three fully connected layers and the ReLU layer are added.

Fig. 1: Synthetic CNN
Fig. 1: Synthetic CNN

Model for Screen Content IQA: Here a model based on neural network for screen content image quality assessment called SCIQA is used. The SCI CNN architecture is shown in figure 2. It consists of 8 convolution layers, 4 max-pooling layers, and 2 fully connected layers. All convolution layers have a filter size of 3 x 3 with stride of 1 pixel. A 2 x 2 pixel kernel with stride of 2 pixels is used in each pooling layer. Each convolutional layer's boundary is padded with zeros to improve network speed.

Fig. 2: SCIQA Model
Fig. 2: SCIQA Model

c) Subjective Score

After the model has been trained, it is used to predict subjective scores for the distorted image. As illustrated in Fig. 2, the trained network is connected to a global average pooling layer before the fully connected layers. A 128-dimensional feature vector is created by averaging the feature map over the spatial domain. The adaptive moment estimation optimizer (ADAM) was used to change the normal stochastic gradient descent approach for better optimization convergence.

VI. EXPERIMENT RESULTS AND ANALYSIS

a) Hardware and Software

The experiments has been conducted, and the results were obtained with a laptop with Intel Processor, 8 GB RAM, and 512 GB SDD. As for software, we have used Python as the programming language, and the libraries such as TensorFlow, Keras, SciPy, Matplotlib, etc. in the Jupyter Notebook. The input pipeline for the model is created using TFDS API.

b) IQA Dataset

The IQA datasets consists of distorted images along with their corresponding pristine images. It also have subjective quality scores for distorted images which is obtained after conducting a psychometric experiments using human subjects. Human opinions are taken for these distorted images with reference to pristine images using some pre-defined range for quality measurement. Various IQA datasets were utilized to measure the performance of the proposed algorithm: LIVE IQA dataset, LIVE multiply distorted (LIVE MD) dataset, and UniMiB MD-IVL dataset. The summary of datasets is given in Table I.

  • The LIVE IQA dataset consists of following types of distortion: WN, JP2K compression, GB, and Rayleigh fast-fading channel distortion [20][21][22].

  • The LIVE MD dataset consists of two categories of images based on distortion combinations applied. First category has images distorted by GB along with JPEG and the second category has images distorted by combination of WN and GB [23].

  • The IVL dataset is generated from 10 reference images which is selected from various samples both in terms of low-level features (frequencies, colors) and high level features [24]. This dataset consists of multiple distorted images with 400 images distorted by noise and JPEG distortions.

Cardinal rating is provided by human observer for all distorted images corresponding to their reference images in the dataset from a pre-defined scale which is considered as Mean Opinion Score (MOS). Hence, each distorted image in the dataset has a corresponding ground-truth subjective quality score.

Table 1: Summary of IQA Datasets Used
DatasetReferencesDistortionTotal Samples
LIVE IQA295982
LIVE MD152450
MD-IVL102400

c) Evaluation Metrics

Unlike traditional pixel-based metrics like PSNR, SSIM, etc. which were used in the past for evaluating IQA algorithms, here the evaluation of the IQA algorithm is done using two statistical measures: SROCC and PLCC i.e., Spearman's rank-order correlation coefficient and Pearson's linear correlation coefficient respectively. The PLCC is calculated using the following formula:

( 1 ) P L C C = i = 1 n ( S ^ i μ S ^ ) ( S i μ S ) i = 1 n ( S ^ i μ S ^ ) 2 ( S i μ S ) 2

where S^i and Si are the predicted and ground-truth subjective scores of the ith image, and μ S^ and μ S denote the mean of each. The SROCC is calculated using the following formula:

( 2 ) S R O C C = 1 6 d i 2 n ( n 2 1 )

where n denotes the number of images and is the difference between predicted score and ground-truth score of image.

d) Results and Analysis

i. Performance on Individual Distortion Types

There are 5 distortion types in LIVE IQA dataset. The distortion types are Fast Fading (FF), JPEG, Gaussian Blur (GB), JP2K, and White Noise (WN). The PLCC and SROCC values for each individual distortion type is evaluated using the DIQA [25] framework. In Table II the PLCC and SROCC values are compared based on the individual distortion type using DIQA framework. For WN, the PLCC and SROCC values are highest whereas for JPEG, it is the lowest. Since JPEG affects the image less compared to other distortion types, so the highest values are for WN distortion type.

Table 2: Comparison of PLCC and SROCC values for different distortion types on LIVE IQA Dataset using DIQA [25] framework In Table II, the PLCC and SROCC values are compared based on the individual distortion type using DNSSCIQ frame-work.
Distortion TypePLCCSROCC
JPEG0.97130.9551
JP2K0.97590.9686
GB0.97670.9713
WN0.98810.9918
FF0.97480.9622
Table 3: Comparison of PLCC and SROCC values for different distortion types on LIVE IQA Dataset using DNSSCIQ framework
Distortion TypePLCCSROCC
JPEG0.98270.9624
JP2K0.96930.9656
GB0.97270.9697
WN0.98810.9918
FF0.94130.9447

Figure 3 shows the comparison of SROCC and PLCC values for various distortion types in the LIVE IQA dataset using DNSSCIQ framework.
Table 4: Comparison of PLCC and SROCC values for different model depth on LIVE IQA Dataset
Model DepthPLCCSROCC
50.96990.9649
60.97690.9712
70.97990.9752
80.98090.9742
90.97670.9738
100.97920.9730

ii. Effect of Model Depth

To determine the influence of model depth, six models with different numbers of convolution layers of DIQA [25] was used.

Convolution layers 1 to 4 and convolution layer 8 was used for the shortest setting. After the Conv6 layer, two 3 × 3 convolution layers with 64 filters were appended in the longest setting. Figure 4 shows the accuracy comparison among the models on the LIVE IQA dataset.

Fig. 4: Comparison of PLCC and SROCC values according to model depth
Fig. 4: Comparison of PLCC and SROCC values according to model depth

Table III shows the PLCC and SROCC values for different model depth. When the depth was 5, the PLCC and SROCC values were the lowest. When the depth is increased, the correlation coefficient got saturated around 0.97. This may cause overfitting when more convolution layers are used. Hence, it is concluded that the 8 convolutional layers are good enough for the proposed framework.

iii. Performance on Individual Datasets

The different datasets are used for evaluating the proposed algorithm. The evaluation metrics such as PLCC and SROCC are used. The datasets are having various types of distortions. In some datasets, various distortion types are combine to produce the distorted image. The DIQA method is evaluated on three different IQA dataset individually. The datasets used are LIVE IQA, LIVE MD and MD IVL. Table V shows the comparison of PLCC and SROCC values for individual datasets using DIQA method. For LIVE IQA dataset, the PLCC and SROCC values are highest

Table 5: Comparison of PLCC and SROCC values for different IQA Datasets using DIQA framework.
DatasetPLCCSROCC
LIVE IQA0.98090.9742
LIVE MD0.95450.9561
MD IVL0.96220.9617
Table 6: Comparison of PLCC and SRCC values for different IQA Datasets using DNSSCIQ
DatasetPLCCSRCC
LIVE IQA0.98670.9799
LIVE MD0.96560.9685
MD IVL0.96960.9702

The PLCC and SROCC values are compared for various IQA datasets like LIVE, LIVE MD and MD IVL in figure 5.

Fig. 5: Comparison of PLCC and SROCC values for various IQA datasets using DNSSCIQ framework iv. Reliability Map
Fig. 5: Comparison of PLCC and SROCC values for various IQA datasets using DNSSCIQ framework iv. Reliability Map

To find the effect of reliability map, the outputs of various configuration is shown in Table VII. It shows that there is an improvement in performance when reliability map is used. Reliability map helps to create homogeneity across the image irrespective of low-frequency components or high-frequency components in the distorted image. This provides the information about the importance of reliability map.

Table 7: Comparison of PLCC and SROCC values with and without Reliability Maps
Reliability MapPLCCSROCC
w/o0.95450.9561
w0.98090.9742

v. NR-IQA Methods

In Table VIII, the PLCC and SROCC metrics of different methods are compared. The different methods are Deep CNN Based Blind Image Quality Predictor (DIQA) [25], Synthetic Convolutional Neural Network (S-CNN) and Screen Content Image Quality Assessment

Table 8: Comparison of PLCC and SROCC values for different method on LIVE IQA Dataset (SCIQA). The S-CNN is having highest PLCC and SROCC values.
MethodPLCCSROCC
DIQA0.98090.9742
S-CNN0.98670.9799
SCIQA0.93380.9229

Figure 6 shows the reference image on the left and distorted image with gaussian blur on the right. The image is obtained from LIVE IQA dataset.

Fig. 3: Comparison of PLCC and SROCC values for various distortion types using DNSSCIQ framework
Fig. 3: Comparison of PLCC and SROCC values for various distortion types using DNSSCIQ framework
Fig. 6: Reference Image on left and Distorted Image (Gaussian Blur) on right
Fig. 6: Reference Image on left and Distorted Image (Gaussian Blur) on right

Figure 7 shows the reference image, distorted image in grayscale, error map, reliability map, perceptual error map and sensitivity map in this order. The image is obtained from LIVE IQA dataset.

Fig. 7: Reference Image, Distorted Image (Gray Scale), Error Map, Reliability Map, Perceptual Error Map, and Sensitivity Map
Fig. 7: Reference Image, Distorted Image (Gray Scale), Error Map, Reliability Map, Perceptual Error Map, and Sensitivity Map

vi. Correlation Plot

Correlation plot shows the correlation between any numerical variables. The correlation coefficient is calculated to determine the correlation between two variables.

Figure 8 shows the correlation plot of ground truth and predicted subjective scores. The ground truth scores are pro-vided in the dataset for each distorted image and DNSSCIQ

A plot to show the correlation between predicted and ground truth subjective score Fig. 8: Correlation Plot framework is used to obtain the predicted subjective score. From the plot, it is concluded that DNSSCIQ is able to calculate the subjective scores almost close to ground-truth values.
Fig. 9: Loss vs. Epoch Graph
Fig. 9: Loss vs. Epoch Graph

vii. Loss Graph

Figure 9 shows loss vs. epoch graph. Here mean squared error is used as loss function. The loss decreases as the number of epochs increases during training. The performance of the model improves with decrease in the loss.

VII. CONCLUSION

A deep CNN-based approach for Non-Screen Content and Screen Content IQA called DNSSCIQ is proposed. In the DNSSCIQ, the input normalization for the distorted images are done first. Then, the distorted image along with its ground-truth subjective score is provided to the neural network for training to obtain more meaningful feature maps. Once the training is completed, the feature maps are globally average pooled and fed the fully connected layers to get the final subjective score of the distorted image. The performance of the DNSSCIQ is good irrespective of the dataset selected is shown by using various datasets from different sources for training and final quality prediction. In addition to this, distortion-specific evaluation of different datasets is done and the output is compared.

References

24 Cites in Article
  1. 1. Yuming Li,Lai-Man Po,Xuyuan Xu,Litong Feng,Fang Yuan,Chun-Ho Cheung,Kwok-Wai Cheung (2015). No-reference image quality assessment with shearlet transform and deep neural networks. Neurocomputing, 154, 94-109.
  2. 2. A Mittal,A Moorthy,A Bovik (2012). Noreference image quality assessment in the spatial domain. IEEE Trans. Image Process, 21(12), 4695-4708.
  3. 3. C Li,A Bovik,X Wu (2011). Blind image quality assessment using a general regression neural network. IEEE Trans. Neural Netw, 22(5), 793-799.
  4. 4. A Moorthy,A Bovik (2011). Blind Image Quality Assessment: From Natural Scene Statistics to Perceptual Quality. IEEE Transactions on Image Processing, 20(12), 3350-3364.
  5. 5. H Tang,N Joshi,A Kapoor (2011). Learning a blind measure of perceptual image quality. Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), 305-312.
  6. 6. J Xu,P Ye,D Doermann (2016). Blind image quality assessment based on high order statistics aggregation. IEEE Transactions on Image Processing, 25(9)
  7. 7. Qiaohong Li,Weisi Lin,Jingtao Xu,Yuming Fang (2016). Blind image quality assessment using statistical structural and luminance features. IEEE Transactions on Multimedia, 18(12)
  8. 8. Jongyoo Kim,Sanghoon Lee (2017). Deep Learning of Human Visual Sensitivity in Image Quality Assessment Framework. 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 1969-1977.
  9. 9. Yuming Li,Lai-Man Po,Xuyuan Xu,Litong Feng,Fang Yuan,Chun-Ho Cheung,Kwok-Wai Cheung (2015). No-reference image quality assessment with shearlet transform and deep neural networks. Neurocomputing, 154, 94-109.
  10. 10. Xialei Liu,Joost Van De Weijer,Andrew Bagdanov (2017). RankIQA: Learning from Rankings for No-Reference Image Quality Assessment. 2017 IEEE International Conference on Computer Vision (ICCV), 1040-1049.
  11. 11. M Saad,A Bovik,C Charrier (2012). Blind Image Quality Assessment: A Natural Scene Statistics Approach in the DCT Domain. IEEE Transactions on Image Processing, 21(8), 3339-3352.
  12. 12. K Ma,W Liu,K Zhang,Z Duanmu,Z Wang,W Zuo (2018). End-to-End Blind Image Quality Assessment Using Deep Neural Networks. IEEE Transactions on Image Processing, 27(3), 1202-1213.
  13. 13. Fei Gao,Yi Wang,Panpeng Li,Min Tan,Jun Yu,Yani Zhu,Deepsim (2017). Deep similarity for image quality assessment. Neurocomputing, 257
  14. 14. Xiongkuo Min,Guangtao Zhai,Ke Gu,Yutao Liu,Xiaokang Yang (2018). Blind Image Quality Estimation via Distortion Aggravation. IEEE Transactions on Broadcasting, 64(2), 508-517.
  15. 15. Hossein Talebi,Peyman Milanfar (2018). NIMA: Neural Image Assessment. IEEE Transactions on Image Processing, 27(8), 3998-4011.
  16. 16. W Hou,X Gao,D Tao,X Li (2015). Blind image quality assessment via deep learning. IEEE Trans. Neural Netw. Learn. Syst, 26(6), 1275-1286.
  17. 17. S Bosse,D Maniry,K U¨ller,T Wiegand,W Samek (2018). Deep Neural Networks for No-Reference and Full-Reference Image Quality Assessment. IEEE Transactions on Image Processing, 27(1), 206-219.
  18. 18. Dounia Hammou,Sid Fezza,Wassim Hamidouche (2021). EGB: Image Quality Assessment based on Ensemble of Gradient Boosting. 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 541-549.
  19. 19. Le Kang,Peng Ye,Yi Li,David Doermann (2015). Simultaneous estimation of image quality and distortion via multi-task convolutional neural networks. 2015 IEEE International Conference on Image Processing (ICIP), 2791-2795.
  20. 20. H Sheikh,Z Wang,L Cormack,A Bovik LIVE Image Quality Assessment Database Release 2. LIVE Image Quality Assessment Database Release 2
  21. 21. H Sheikh,M Sabir,A Bovik (2006). A Statistical Evaluation of Recent Full Reference Image Quality Assessment Algorithms. IEEE Transactions on Image Processing, 15(11), 3440-3451.
  22. 22. Z Wang,A Bovik,H Sheikh,E Simoncelli (2004). Image quality assessment: from error visibility to structural similarity. IEEE Transactions on Image Processing, 13(4), 600-612.
  23. 23. Dinesh Jayaraman,Anish Mittal,Anush Moorthy,Alan Bovik (2012). Objective quality assessment of multiply distorted images. 2012 Conference Record of the Forty Sixth Asilomar Conference on Signals, Systems and Computers (ASILOMAR), 1693-1697.
  24. 24. J Kim,A Nguyen,S Lee (2019). Deep CNN-Based Blind Image Quality Predictor. IEEE Transactions on Neural Networks and Learning Systems, 30(1), 11-24.

Funding

No external funding was declared for this work.

Conflict of Interest

The authors declare no conflict of interest.

Ethical Approval

No ethics committee approval was required for this article type.

Data Availability

Not applicable for this article.

How to Cite This Article

Dilip Chaudhary. 2026. "Deep CNN Model for Non-Screen Content and Screen Content Image Quality Assessment". Global Journal of Computer Science and Technology - D: Neural & AI GJCST-D Volume 22 (GJCST Volume 22 Issue D1).

Download Citation

Advanced deep CNN model for non-screen content quality assessment.
Journal Specifications

Crossref Journal DOI 10.17406/gjcst

Print ISSN 0975-4350

e-ISSN 0975-4172

Keywords
Classification
GJCST-D Classification F.1.1
Version of record

v1.2

Issue date
January 22, 2022

Language
English
Order Article Reprint
Experiance in AR

Explore published articles in an immersive Augmented Reality environment. Our platform converts research papers into interactive 3D books, allowing readers to view and interact with content using AR and VR compatible devices.

Read in 3D

Your published article is automatically converted into a realistic 3D book. Flip through pages and read research papers in a more engaging and interactive format.

Article Matrices
Total Views: 1.1K
Total Downloads: 90
All Trends

Request Access

Please fill out the form below to request access to this research paper. Your request will be reviewed by the editorial or author team.
X

This is the heading

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Ut elit tellus, luctus nec ullamcorper mattis, pulvinar dapibus leo.

High-quality academic research articles on global topics and journals.

Deep CNN Model for Non-Screen Content and Screen Content Image Quality Assessment

Dilip Chaudhary
Dilip Chaudhary
Venkatesh
Venkatesh