Search arXiv⌕ Search

arXiv · 2009.03411

Deep Local and Global Spatiotemporal Feature Aggregation for Blind Video Quality Assessment

Abstract

In recent years, deep learning has achieved promising success for multimedia quality assessment, especially for image quality assessment (IQA). However, since there exist more complex temporal characteristics in videos, very little work has been done on video quality assessment (VQA) by exploiting powerful deep convolutional neural networks (DCNNs). In this paper, we propose an efficient VQA method named Deep SpatioTemporal video Quality assessor (DeepSTQ) to predict the perceptual quality of various distorted videos in a no-reference manner. In the proposed DeepSTQ, we first extract local and global spatiotemporal features by pre-trained deep learning models without fine-tuning or training from scratch. The composited features consider distorted video frames as well as frame difference maps from both global and local views. Then, the feature aggregation is conducted by the regression model to predict the perceptual video quality. Finally, experimental results demonstrate that our proposed DeepSTQ outperforms state-of-the-art quality assessment algorithms.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Wei Zhou, Zhibo Chen. 2020-09-07. Deep Local and Global Spatiotemporal Feature Aggregation for Blind Video Quality Assessment. https://arxiv.org/abs/2009.03411

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Image Reconstruction from Phase with Untrained Neural Priors

Fourier phase encodes important spatial image structure, but recovering an image without measured spectral magnitude requires additional constraints and leaves absolute intensity ambiguous. We propose a projection-based two-stage framework that combines Fourier-phase and spatial-support constraints with an image-specific neural prior. The first stage alternates constraint enforcement with regularized neural-prior updates, while the second performs phase/support refinement alone with guaranteed convergence. We evaluate two neural-prior implementations on the same 77 microscopy images and compare them with a constraint-only baseline. After 500 final refinement passes, the best-performing variant achieves 31.41 dB pooled PSNR, 35.75 dB mean PSNR, and 0.9531 mean SSIM, improving pooled PSNR by 1.51~dB and reducing pooled MSE by 29.3% relative to the baseline. The results demonstrate the benefit of combining neural guidance with explicit constraint refinement at the evaluated iteration budget, while showing that lower phase residual alone does not guarantee greater reconstruction accuracy.

eess.IV↗

Deep Pseudo-Proximal Map: A Self-Supervised Data-Fitting Agent for Iterative Reconstruction

Iterative algorithms for inverse imaging problems split reconstruction into alternating subproblems, one of which requires evaluating the data-fitting proximal map. In many practical cases, this proximal map has no analytic solution and so must be approximated by an inner loop solver at every iteration, compounding the cost of the outer reconstruction loop. To address this, we propose the pseudo-proximal map (PPM), a reformulation of the data-fitting proximal map as the minimum mean square error estimate of a synthetic probabilistic model. We implement the deep PPM as a self-supervised neural network trained only on sampled Gaussian noise, requiring no ground-truth training images. The deep PPM can be trained for any operator for which the forward model $A$ and its transpose $A^T$ can be evaluated, with provable equivalence to the proximal map when $A$ is linear. We validate the deep PPM on Gaussian deblurring and 4x super-resolution, where the proximal map has an analytic solution, and on X-ray computed tomography (XCT), where no analytic solution exists. Used as the data-fitting agent within an iterative reconstruction method, the deep PPM reproduces the reference reconstruction to within 1% NRMSE for all three operators, and for XCT, it replaces the inner conjugate-gradient loop with a single network evaluation that is 18x faster.

eess.IV↗

Bridging the Inter-Domain Gap through Low-Level Features for Cross-Modal Medical Image Segmentation

This paper addresses cross-modal medical image segmentation, focusing on MRI-CT transfer in a source-only domain generalization setting. During training, only source-modality samples are available, while unlabeled target-modality images are used for testing. We propose LowBridge, which builds on the observation that cross-modal images share similar low-level features (e.g. edges) as they depict the same types of anatomical structures. Specifically, we first train a generative model to recover the source images from their edge features, followed by training a segmentation model on the generated source images, separately. At test time, edge features from the target images are input to the pretrained generative model to generate source-style target domain images, which are then segmented using the pretrained segmentation network. Experiments on various public datasets demonstrate that LowBridge achieves state-of-the-art performance, outperforming ten existing approaches. Ablation studies further show that LowBridge is compatible with different types of generative and segmentation models, suggesting its generalizability and potential to benefit from future advances in these models. The code will be available at https://github.com/JoshuaLPF/LowBridge.

eess.IV↗