You will be provided with a reference and some statements. Please determine whether each statement is 'supported', 'unsupported', or 'unknown' with respect to the reference. Please note:
First, assess whether the reference contains any valid content. If the reference contains no valid information, such as a 'page not found' message, then all statements should be considered 'unknown'.
If the reference is valid, for a given statement: if the facts or data it contains can be found entirely or partially within the reference, it is considered 'supported' (data accepts rounding); if all facts and data in the statement cannot be found in the reference, it is considered 'unsupported'.

You should return the result in a JSON list format, where each item in the list contains the statement's index and the judgment result, for example:
[
    {
        "idx": 1,
        "result": "supported"
    },
    {
        "idx": 2,
        "result": "unsupported"
    }
]

Below are the reference and statements:
<reference>
[PDF] A Survey of Video-based Action Quality Assessment | Semantic Scholar

This paper provides a comprehensive survey of existing papers on video-based action quality assessment and summarized the methods of sports and medical careaccording to the model categories and publishing institutions according to the characteristics of the two fields. Human action recognition and analysis have great demand and important application significance in video surveillance, video retrieval, and human-computer interaction. The task of human action quality evaluation requires the intelligent system to automatically and objectively evaluate the action completed by the human. The action quality assessment model can reduce the human and material resources spent in action evaluation and reduce subjectivity. In this paper, we provide a comprehensive survey of existing papers on video-based action quality assessment. Different from human action recognition, the application scenario of action quality assessment is relatively narrow. Most of the existing work focuses on sports and medical care. We first introduce the definition and challenges of human action quality assessment. Then we present the existing datasets and evaluation metrics. In addition, we summarized the methods of sports and medical care according to the model categories and publishing institutions according to the characteristics of the two fields. At the end, combined with recent work, the promising development direction in action quality assessment is discussed.

Opens in a new window

Opens an external website

Opens an external website in a new window

This website utilizes technologies such as cookies to enable essential site functionality, as well as for analytics, personalization, and targeted advertising.

Privacy Policy

Accept

Deny Non-Essential

Manage Preferences

Skip to search form
Skip to main content
Skip to account menu
Search 238,063,513 papers from all fields of science
Search
Sign In
Create Free Account
DOI:
10.1109/INSAI54028.2021.00029
Corpus ID: 248266561
A Survey of Video-based Action Quality Assessment
@article{Wang2021ASO,
title={A Survey of Video-based Action Quality Assessment},
author={Shunli Wang and Dingkang Yang and Peng Zhai and Qing Yu and Tao Suo and Zhan Sun and Ka Li and Lihua Zhang},
journal={2021 International Conference on Networking Systems of AI (INSAI)},
year={2021},
pages={1-9},
url={https://api.semanticscholar.org/CorpusID:248266561}
}
Shunli Wang
,
Dingkang Yang
,
+5 authors

Lihua Zhang
Published
in
International Conference on…

1 November 2021
Computer Science
2021 International Conference on Networking Systems of AI (INSAI)
TLDR
This paper provides a comprehensive survey of existing papers on video-based action quality assessment and summarized the methods of sports and medical careaccording to the model categories and publishing institutions according to the characteristics of the two fields.
Expand
[PDF] Semantic Reader
Save to Library
Save
Create Alert
Alert
Cite
Share
27 Citations
Highly Influential Citations
2
Background Citations
7
Methods Citations
1
View All
Figures and Tables
27 Citations
66 References
Related Papers
Figures and Tables from this paper
figure 1
figure 2
figure 3
figure 4
figure 5
figure 6
table I
table II
View All 8 Figures & Tables
27 Citations
Citation Type
Has PDF
Author
More Filters
More Filters
Filters
Sort by Relevance
Sort by Most Influenced Papers
Sort by Citation Count
Sort by Recency
Expert Systems With Applications
Jiang Liu
Hua-Sheng Wang
Katarzyna Stawarz
Shiyin Li
Yao Fu
Han-Tao Liu
Computer Science
TLDR
This review presents an overview of various aspects of AQA, including existing applications, data acquisition methods, public datasets, state-of-the-art methods and evaluation metrics, which can be a helpful guide for researchers to explore AQA.
Expand
2 Excerpts
Save
A comprehensive survey of action quality assessment: Method and benchmark
Kang-Lei Zhou
Ruizhi Cai
Liyuan Wang
Hubert P. H. Shum
Xiaohui Liang
Computer Science, Medicine
Pattern Recognition
2026
TLDR
A modality-driven hierarchical taxonomy is proposed that organizes existing methods into video-based, skeleton-based, and multi-modal approaches, and analyzes the methodological evolution of representative models to establish a unified benchmark for representative video-based AQA methods.
Expand
28
PDF
Save
Hierarchical Graph Convolutional Networks for Action Quality Assessment
Kang-Lei Zhou
Yue Ma
Hubert P. H. Shum
Xiaohui Liang
Computer Science
IEEE transactions on circuits and systems for…
2023
TLDR
A hierarchical graph convolutional network (GCN) is proposed that corrects semantic information confusion through clip refinement, generating the ‘shot’ as the basic action unit and constructs a scene graph by combining several consecutive shots into meaningful scenes to capture local dynamics.
Expand
76
PDF
1 Excerpt
Save
A multimodal framework for action quality assessment of continuous sit-ups
Meng Tian
Peirui Bai
+4 authors

Meng-Hao Xu
Computer Science
International Conference on Robotics and Sensor…
2026
TLDR
A multi-modal framework is proposed to adaptively determine action clips and highlight key segments based on Transformer framework, which demonstrated that the round segments of continuous sit-ups can be clipped accurately, and the key frames can be extracted with better interpretability.
Expand
2 Excerpts
Save
FineSkiing: A Fine-grained Benchmark for Skiing Action Quality Assessment
Yongji Zhang
Siqi Li
Yue Gao
Yu Jiang
Computer Science
arXiv.org
2025
TLDR
This paper constructs the first AQA dataset containing fine-grained sub-score and deduction annotations for aerial skiing, which will be released as a new benchmark and proposes a novel AQA method, named JudgeMind, which significantly enhances performance and reliability by simulating the judgment and scoring mindset of professional referees.
Expand
1
[PDF]
1 Excerpt
Save
Adaptive Spatiotemporal Graph Transformer Network for Action Quality Assessment
Jiang Liu
Hua-Sheng Wang
+4 authors

Han-Tao Liu
Computer Science
IEEE transactions on circuits and systems for…
2025
TLDR
An adaptive spatiotemporal graph transformer network (ASGTN) that combines multiple graph structures and transformer attention mechanisms to capture both local and global contextual information within and across clips in a long video is proposed.
Expand
27
PDF
1 Excerpt
Save
A Decade of Action Quality Assessment: Largest Systematic Survey of Trends, Challenges, and Future Directions
Hao Yin
Paritosh Parmar
Daoliang Xu
Yang Zhang
Tian-You Zheng
Weiwei Fu
Computer Science
International Journal of Computer Vision
2026
TLDR
A thorough survey of the AQA landscape is presented, systematically reviewing over 200 research papers using the preferred reporting items for systematic reviews and meta-analyses (PRISMA) framework, covering foundational concepts and definitions, then move to general frameworks and performance metrics, and finally discuss the latest advances in methodologies and datasets.
Expand
19
Highly Influenced
PDF
6 Excerpts
Save
CPR-Coach: Recognizing Composite Error Actions Based on Single-Class Training
Shunli Wang
Qing Yu
+7 authors

Lihua Zhang
Medicine, Computer Science
Computer Vision and Pattern Recognition
2024
TLDR
A human-cognition-inspired framework named ImagineNet is proposed to improve the model's multi-error recognition performance under restricted supervision to solve the unavoidable “Single-class Training & Multi-class Testing” problem.
Expand
9
[PDF]
3 Excerpts
Save
SkillSpotter: Pose-Aware Multi-View Skilled Action Detection and Grading in Ego-Exo Videos
Björn Braun
Christian Holz
Computer Science
arXiv.org
2026
TLDR
SkillSpotter is introduced, a pose-aware multi-view architecture that jointly detects and grades skilled actions through three task-specific modules that transfer to other temporal action detection models with consistent gains, and generalizes beyond Ego-Exo4D to HoloAssist.
Expand
[PDF]
2 Excerpts
Save
PE3DNet: A “Pull-Up” Action Quality Assessment Network Based on Fusion of RGB Image Data and Key Point Optical Flow
Rui‐Zhe Yang
Xiaona Xiu
Jian Wang
Ruyao Wang
Computer Science
2024 7th International Conference on Pattern…
2024
TLDR
An action quality assessment network PE3DNet is proposed for “Pull-up” video streams, which uses OpenPifPaf to generate key points, constructs key points optical flow data using key points, and adopts P3D residual blocks to construct a dual-stream network that fuses RGB image data and key point optical flow to achieve the action quality assessment of “Pull-up”.
Expand
2
Save
...
1
2
3
...
66 References
Citation Type
Has PDF
Author
More Filters
More Filters
Filters
Sort by Relevance
Sort by Most Influential Papers
Sort by Citation Count
Sort by Recency
Assessing the Quality of Actions
H. Pirsiavash
Carl Vondrick
A. Torralba
Computer Science
European Conference on Computer Vision
2014
TLDR
A learning-based framework that takes steps towards assessing how well people perform actions in videos by training a regression model from spatiotemporal pose features to scores obtained from expert judges and can provide interpretable feedback on how people can improve their action.
Expand
288
Highly Influential
PDF
3 Excerpts
Save
TSA-Net: Tube Self-Attention Network for Action Quality Assessment
Shunli Wang
Dingkang Yang
Peng Zhai
Chixiao Chen
Lihua Zhang
Computer Science
ACM Multimedia
2021
TLDR
A Tube Self-Attention Network (TSA-Net) is proposed for action quality assessment (AQA), which can efficiently generate rich spatio-temporal contextual information by adopting sparse feature interactions and achieves the Spearman's Rank Correlation results.
Expand
108
Highly Influential
[PDF]
7 Excerpts
Save
End-To-End Learning for Action Quality Assessment
Yongjun Li
Xiujuan Chai
Xilin Chen
Computer Science
Pacific Rim Conference on Multimedia
2018
TLDR
An end-to-end framework is proposed based on fragment-based 3D convolutional neural network to realize the action quality assessment in videos to narrow the gap between the predictions and ground-truth scores as well as making the predictions satisfy the ranking constraint.
Expand
58
1 Excerpt
Save
Uncertainty-Aware Score Distribution Learning for Action Quality Assessment
Yansong Tang
Zanlin Ni
+4 authors

Jie Zhou
Computer Science
Computer Vision and Pattern Recognition
2020
TLDR
An uncertainty-aware score distribution learning (USDL) approach for action quality assessment (AQA) is proposed, which regards an action as an instance associated with a score distribution, which describes the probability of different evaluated scores.
Expand
200
[PDF]
1 Excerpt
Save
Action Assessment by Joint Relation Graphs
Jiahui Pan
Jibin Gao
Weishi Zheng
Computer Science
IEEE International Conference on Computer Vision
2019
TLDR
This work presents a new model to assess the performance of actions from videos, through graph-based joint relation modelling, and proposes two novel modules, the Joint Commonality Module and the Joint Difference Module, for joint motion learning.
Expand
156
PDF
2 Excerpts
Save
What and How Well You Performed? A Multitask Learning Approach to Action Quality Assessment
Paritosh Parmar
B. Morris
Computer Science
Computer Vision and Pattern Recognition
2019
TLDR
This paper proposes to learn spatio-temporal features that explain three related tasks - fine-grained action recognition, commentary generation, and estimating the AQA score, and shows that the MTL approach outperforms STL approach using two different kinds of architectures: C3D-AVG and MSCADC.
Expand
240
[PDF]
2 Excerpts
Save
Temporal Segment Networks for Action Recognition in Videos
Limin Wang
Yuanjun Xiong
+4 authors

L. van Gool
Computer Science
IEEE Transactions on Pattern Analysis and Machine…
2019
TLDR
The proposed TSN framework, called temporal segment network (TSN), aims to model long-range temporal structure with a new segment-based sampling and aggregation scheme and won the video classification track at the ActivityNet challenge 2016 among 24 teams.
Expand
983
[PDF]
1 Excerpt
Save
Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset
João Carreira
Andrew Zisserman
Computer Science
Computer Vision and Pattern Recognition
2017
TLDR
I3D models considerably improve upon the state-of-the-art in action classification, reaching 80.2% on HMDB-51 and 97.9% on UCF-101 after pre-training on Kinetics, and a new Two-Stream Inflated 3D Conv net that is based on 2D ConvNet inflation is introduced.
Expand
9,945
[PDF]
1 Excerpt
Save
Graph-based analysis of physical exercise actions
Oya Celiktutan
C. B. Akgül
Christian Wolf
B. Sankur
Computer Science
MIIRH '13
2013
TLDR
A graph-based method to align two dynamic skeleton sequences is developed, and applied to both action recognition tasks as well as to the objective quantification of the goodness of the action performance.
Expand
38
PDF
1 Excerpt
Save
Manipulation-Skill Assessment from Videos with Spatial Attention Network
Zhenqiang Li
Yifei Huang
Minjie Cai
Yoichi Sato
Computer Science
2019 IEEE/CVF International Conference on…
2019
TLDR
A novel RNN-based spatial attention model is proposed that considers accumulated attention state from previous frames as well as high-level information about the progress of an undergoing task in automatic skill assessment.
Expand
66
[PDF]
3 Excerpts
Save
...
1
2
3
4
5
...
Related Papers
Show More
2/0
Related Papers
Stay Connected With Semantic Scholar
Sign Up
What Is Semantic Scholar?
Semantic Scholar is a free, AI-powered research tool for scientific literature, based at Ai2.
Learn More
About
About Us
Publishers
Blog
(opens in a new tab)
Ai2 Careers
(opens in a new tab)
Product
Product Overview
Semantic Reader
Scholar's Hub
Beta Program
Release Notes
API
API Overview
API Tutorials
API Documentation
(opens in a new tab)
API Gallery
Research
Publications
(opens in a new tab)
Research Careers
(opens in a new tab)
Resources
(opens in a new tab)
Help
FAQ
Librarians
Tutorials
Contact
Proudly built by
Ai2
(opens in a new tab)
Collaborators & Attributions
•
Terms of Service
(opens in a new tab)
•
Privacy Policy
(opens in a new tab)
•
API License Agreement
The Allen Institute for AI
(opens in a new tab)
</reference>

<statements>
1. While early AQA implementations relied on single-score linear regression trained via mean squared error, modern frameworks employ pairwise deep ranking and uncertainty-aware score distribution learning to accommodate natural judging variations.
2. The primary performance baseline for MTL-AQA is 0.90–0.94 for multi-task spatial-temporal CNNs.
</statements>

Begin the assessment now. Output only the JSON list, without any conversational text or explanations.