You will be provided with a reference and some statements. Please determine whether each statement is 'supported', 'unsupported', or 'unknown' with respect to the reference. Please note:
First, assess whether the reference contains any valid content. If the reference contains no valid information, such as a 'page not found' message, then all statements should be considered 'unknown'.
If the reference is valid, for a given statement: if the facts or data it contains can be found entirely or partially within the reference, it is considered 'supported' (data accepts rounding); if all facts and data in the statement cannot be found in the reference, it is considered 'unsupported'.

You should return the result in a JSON list format, where each item in the list contains the statement's index and the judgment result, for example:
[
    {
        "idx": 1,
        "result": "supported"
    },
    {
        "idx": 2,
        "result": "unsupported"
    }
]

Below are the reference and statements:
<reference>
[2405.15813] From CNNs to Transformers in Multimodal Human
Action Recognition: A Survey



From CNNs to Transformers in Multimodal Human
Action Recognition: A Survey
DOI:
X.X
Conference:
Make sure to enter the correct
conference title from your rights confirmation email; ; NY
Price:
15.00
ISBN:
978-1-4503-XXXX-X/18/06
CCS:
General and reference Surveys and overviews
CCS:
Computing methodologies Activity recognition and understanding
CCS:
Computing methodologies Scene understanding
CCS:
Computing methodologies Neural networks

Muhammad Bilal Shaikh

email:
m.shaikh@ecu.edu.au

OrcID:
0000-0001-9042-5018

Affiliation:
Edith Cowan University

,
270 Joondalup Drive

,
Perth

,
Western Australia

,
Australia

,
6027

,

Douglas Chai

OrcID:
0000-0002-9004-7608

Affiliation:
Edith Cowan University

,
270 Joondalup Drive

,
Perth

,
Western Australia

,
Australia

,
6027

,

Syed Mohammed Shamsul Islam

OrcID:
0000-0002-3200-2903

email:
d.chai@ecu.edu.au

Affiliation:
Edith Cowan University

,
270 Joondalup Drive

,
Perth

,
Western Australia

,
Australia

,
6027

and

Naveed Akhtar

OrcID:
0000-0003-3406-673X

Affiliation:
School of Computing and Information Systems, The University of Melbourne

,
700 Swanston Street

,
Carlton

,
Victoria

,
Australia

,
3010

email:
naveed.akhtar@uwa.edu.au

2024© , 2024;

Abstract.

Due to its widespread applications, human action recognition is one of the most widely studied research problems in Computer Vision. Recent studies have shown that addressing it using multimodal data leads to superior performance as compared to relying on a single data modality. During the adoption of deep learning for visual modelling in the last decade, action recognition approaches have mainly relied on Convolutional Neural Networks (CNNs). However, the recent rise of Transformers in visual modelling is now also causing a paradigm shift for the action recognition task. This survey captures this transition while focusing on Multimodal Human Action Recognition (MHAR). Unique to the induction of multimodal computational models is the process of ‘fusing’ the features of the individual data modalities. Hence, we specifically focus on the fusion design aspects of the MHAR approaches. We analyze the classic and emerging techniques in this regard, while also highlighting the popular trends in the adaption of CNN and Transformer building blocks for the overall problem. In particular, we emphasize on recent design choices that have led to more efficient MHAR models. Unlike existing reviews, which discuss Human Action Recognition from a broad perspective, this survey is specifically aimed at pushing the boundaries of MHAR research by identifying promising architectural and fusion design choices to train practicable models. We also provide an outlook of the multimodal datasets from their scale and evaluation viewpoint. Finally, building on the reviewed literature, we discuss the challenges and future avenues for MHAR.

Keywords:
Multimodal, action recognition, fusion, deep learning, neural networks

1.
Introduction

Figure 1.
(
Left
) Number of relevant publications in recent years, identified with the data collected from the Web of Science. (
Center
) Categories of publication contributions to different sub-fields of Science - generated with data from the Web of Science. (
Right
) Distribution as document type (data collected from Scopus).

Modality broadly refers to the mode in which information is perceived. Humans are able to see, feel, hear and so on, using their senses. Our ability to intelligently process this multimodal information is what makes us highly effective beings. Currently, the field of Artificial Intelligence (AI) is generally concerned with developing capabilities for individual data modalities that encode specific information about our surroundings. In this regard, in the form of recent breakthroughs such as GPT-4
(
OpenAI 2023
)
for multimodal
data (image and text), ChatGPT for text, DALL-E 2
(
Ramesh et al. 2021
)
and Stable Diffusion
(
Rombach et al. 2022
)
for the visual data, we are already witnessing human-level intelligence of machines. However, the ultimate objective to achieve Artificial General Intelligence (AGI) demands much more. The ability to handle and leverage information encoded in multimodal data is one of the basic needs of AGI. With the availability of increasingly powerful computational resources, this realization is fast re-writing the guiding principles for the research communities in different sub-fields of AI.

In the domain of Computer Vision, human action recognition is a long-standing problem
(
Wang et al. 2019c
)
. This article follows Herath et al.
(
Herath et al. 2017
)
, treating human action as “the most elementary human-surrounding interaction with a meaning.” Human action recognition is thus the automated labelling process of human actions within a given sequence of visual frames. Due to its widespread practical applications, ranging from safety
(
Sun and Chen 2022
)
to healthcare
(
Zlatintsi et al. 2020
)
, particularly for COVID
(
Rahman et al. 2021
)
and many other downstream tasks
(
Fassold and Takacs 2019
;
Qi et al. 2019
;
Feng et al. 2017
)
, human action recognition has historically received significant attention within the vision research community. Recent years have also seen the emergence of numerous multimodal approaches for this task, which define the research direction of Multimodal Human Action Recognition (MHAR). Multimodal data fusion refers
to the mixing of features from data of various modalities. Fusion analysis has been extensively used in several domains including activity recognition
(
Min et al. 2016
)
, 3D shape classification
(
Nie et al. 2020
)
and predicting human eye fixation
(
Huang et al. 2020
)
, as evident through relevant works.

Presently, MHAR systems have been explored in various real-world scenarios of indoor and outdoor settings that relate to the application domains of healthcare
(
Bruce et al. 2020
;
Zahin et al. 2019
)
, smart homes
(
Sharif et al. 2020
;
Nan et al. 2019
;
Liciotti et al. 2020
)
, surveillance systems
(
Ullah et al. 2019
;
Khan et al. 2020
;
Batchuluun et al. 2019
)
, social relations recognition
(
Xu et al. 2021
)
and numerous other domains
(
Susan et al. 2019
;
Tingting et al. 2019
)
. It is evident from the existing literature that owing to the diversity of data, the research direction of MHAR has its own unique challenges and opportunities when compared to conventional action recognition tasks
(
Mohsen et al. 2022
)
. To contextualize this, early research in human action recognition was limited to only analysing still color images or videos
(
Guo and Lai 2014
;
Zhu et al. 2016
)
. The MHAR problem demands accounting for different data modalities while providing the opportunity for improved performance due to the availability of complementary information in different modalities. Nevertheless, it is worth noting here that the typical challenges of classical action recognition, (e.g., background clutter, partial object occlusion, viewpoint variations, lighting changes, and execution rate etc.) remain equally pertinent to MHAR techniques.

This survey focuses on MHAR approaches that use computer vision and signal-processing techniques to recognize human actions. Generally, techniques based on multimodal data are motivated based on the expectation to achieve better performance than unimodal data techniques. However, MHAR research has also shown other motivational sources in the form of the development of cost-effective sensing (e.g., ASUS Xtion
(
Inc 2020
)
, Microsoft Kinect
(
Yang et al. 2015
)
, and Intel RealSense
(
Carfagni et al. 2019
)
) and widespread gain in the computational power. Combined, these factors currently propel MHAR research to an unprecedented pace.
Accordingly, there has been a growing number of MHAR-related publications in the last few years - see Fig.
1
(left). Interestingly, these publications have contributed to diverse sub-fields of Science as their application domains - Fig.
1
(middle).
In Fig
1
, we also show the Scopus article-type distribution against the query “multimodal action recognition" to provide an overview of the publication trends in this highly active research direction.

Due to the emerging popularity of the topic, several survey articles have examined MHAR from different perspectives
(
Atrey et al. 2010
;
Wang 2021
;
Zhang et al. 2019a
;
Ahmad and Conci 2019
;
Wang et al. 2018
;
Aggarwal and Xia 2014
;
Shaikh and Chai 2021
;
Han et al. 2017
;
Zhang et al. 2016
;
Zhu et al. 2016
;
Zhang et al. 2019b
;
Chen et al. 2017
;
Zhang et al. 2017
;
Sun et al. 2022
)
.
However, this article is distinct from all previous studies, as we specifically focus and extensively cover the area of multimodal fusion and learning in action recognition. To the best of our knowledge, no prior detailed review article has thoroughly addressed this research direction.
The specific contributions of this survey are summarized as follows.

•

To the best of our knowledge, this is the first comprehensive review of MHAR methods from the perspective of multimodality, with a focus on CNNs and transformer-based techniques.

•

It provides a thorough review of fusion approaches beyond the typical early and late splits used in the implementation of multimodal action recognition.

•

It provides a comprehensive comparison of the existing methods and their performance on benchmark datasets, along with insightful discussions.

•

It pays special attention to the more recent approaches for MHAR, thereby providing readers with an accessible overview of the current state-of-the-art in MHAR.

Figure 2.
A typical multimodal fusion-based action recognition pipeline.

The article is organized as follows: Section
2
discusses multimodal learning methods, covering the original CNN and Transformer architectures and their derivatives. Section
3
summarizes standard benchmark datasets and their state-of-the-art. Fusion mechanisms in multimodal action recognition are explored in Section
4
, highlighting recent advancements in various fusion strategies. Section
5
addresses the challenges in current multimodal action recognition. Section
6
concludes the article.

2.
Multimodal Learning Methods

2.1.
Multimodal Learning

Before discussing multimodal learning, we must first understand unimodal learning, where a neural network processes a single data type.

Given a dataset
T
=
{
x
1
,
…
,
x
n
,
y
1
,
…
,
y
n
}
T=\{{x_{1},...,x_{n},y_{1},...,y_{n}}\}
,
x
i
x_{i}
is the i-th training example and
y
i
y_{i}
is the true label of
x
i
x_{i}
. Training on a single modality, say RGB images, can be formalized by the equation:

(1)

L
⁡
(
C
⁡
(
ϕ
m
​
(
X
)
)
,
y
)
L(C(\phi_{m}(X)),y)

where
ϕ
m
\phi_{m}
is a deep neural network tailored for that modality (like a CNN for images). Its parameters are represented by
⊖
m
\ominus_{m}
.
C
C
is a classifier that predicts the label based on the features extracted by
ϕ
m
\phi_{m}
. This classifier typically uses one or more fully connected layers, and its parameters are denoted by
⊖
c
\ominus_{c}
.

When dealing with real-world problems, often a single data modality is not enough. For example, in video understanding tasks, both video frames (RGB) and audio can provide valuable information. By leveraging multiple modalities, a model can potentially achieve better performance than relying on a single one. A typical multimodal fusion-based system would follow a standard workflow (see Fig.
2
).

Training with multiple modalities can be formalized as:

(2)

L
multi
=
L
⁡
(
C
⁡
(
ϕ
audio
⊕
ϕ
video
)
,
y
)
L_{\textrm{multi}}=L(C(\phi_{\textrm{audio}}\oplus\phi_{\textrm{video}}),y)

here, each modality, audio and video, has its respective deep neural network,
ϕ
audio
\phi_{\textrm{audio}}
and
ϕ
video
\phi_{\textrm{video}}
, designed to extract relevant features. The fusion operation, represented by
⊕
\oplus
, can be a simple concatenation, or more sophisticated operations like weighted sums or attention mechanisms.

2.2.
Learning Methods in CNNs

2.2.1.
Convolution Neural Networks (CNNs)

CNN architectures found in the literature are diverse while sharing common elements. Primarily, they incorporate convolutional layers, where multiple kernels compute various feature maps based on input data. Generating a new feature map involves convolving the input with a learned kernel and applying an activation function to the results. Shared kernels, applied across all spatial input locations, produce individual feature maps, simplifying model complexity and improving trainability. The feature value at a specific location
(
i
,
j
)
(i,j)
in the
k
k
-th feature map of the
l
l
-th layer, denoted as
z
i
,
j
,
k
l
z^{l}_{i,j,k}
, is given by the equation:

(3)

z
i
,
j
,
k
l
=
w
k
l
T
​
x
i
,
j
l
+
b
k
l
z^{l}_{i,j,k}={w^{l}_{k}}^{T}x^{l}_{i,j}+b^{l}_{k}

Here, the shared kernel
w
k
l
w^{l}_{k}
generates the feature map
z
l
:
,
:
,
k
z^{l}_{:,:,k}
. The activation function
a
⁡
(
⋅
)
a(\cdot)
introduces crucial non-linearities, determining the convolutional feature activation value
a
i
,
j
,
k
l
a^{l}_{i,j,k}
as follows:

(4)

a
i
,
j
,
k
l
=
a
⁡
(
z
i
,
j
,
k
l
)
a^{l}_{i,j,k}=a(z^{l}_{i,j,k})

The pooling layer enhances shift-invariance by reducing feature map resolution. For each feature map
a
l
:
,
:
,
k
a^{l}_{:,:,k}
, the pooling function, denoted as pool
(
⋅
)
(\cdot)
, is applied as follows:

(5)

y
i
,
j
,
k
l
=
pool
​
(
a
m
,
n
,
k
l
)
,
∀
(
m
,
n
)
∈
ℝ
i
​
j
y^{l}_{i,j,k}=\text{pool}(a^{l}_{m,n,k}),\forall(m,n)\in\mathbb{R}_{ij}

After multiple convolutional and pooling layers, fully connected layers may exist for high-level reasoning. The output layer, often employing the softmax operator for classification, serves as the final layer. CNN parameters, denoted by
θ
\theta
, including weight vectors and bias terms, are optimized by minimizing a task-specific loss function. For
N
N
input-output pairs
(
x
(
n
)
,
y
(
n
)
)
;
n
∈
[
1
,
N
]
{(x^{(n)},y^{(n)});n\in[1,N]}
, the CNN loss is calculated as:

(6)

ℒ
multi
=
1
N
​
∑
n
=
1
N
λ
⁡
(
θ
,
y
(
n
)
,
o
(
n
)
)
\mathcal{L}_{\textrm{multi}}=\frac{1}{N}\sum_{n=1}^{N}\lambda(\theta;y^{(n)},o^{(n)})

2.2.2.
CNN-based Learning Methods

Table 1.
Summary of CNN-based MHAR methods applied on standard benchmark public datasets using fusion methods. Abbreviations: Modality (S: Skeleton, IR: Infrared, P: Pose, D: Depth), Datasets (U: UTD-MHAD, UT: UT-Kinect, N1:NTU RGB+D 60, N2:NTU RGB+D 120, NW: NW-UCLA, T: Toyota SH, M: MSAR Daily, UCF: UCF51).

Ref.

Modality

Fusion

U

UT

N1

N2

NW

T

M

UCF

(
Islam and Iqbal 2020
)

RGB+S

Middle

95.12

97.56

-

-

-

-

-

-

(
Memmesheimer et al. 2020
)

S+IR

Early

93.3

-

70.8

78.3

-

-

-

-

(
Das et al. 2020
)

RGB+P

Middle

-

-

95.5

93.5

86.3

60.8

-

-

(
De Boissiere and Noumeir 2020
)

P+IR

Late

-

-

91.6

-

-

-

-

-

(
Vaezi Joze et al. 2020
)

RGB+P

Middle

-

-

91.99

-

-

-

-

-

(
Perez-Rua et al. 2019
)

RGB+P

Hybrid

-

-

90.04

-

-

-

-

-

(
Das et al. 2019
)

RGB+P

Late

-

-

92.2

-

90.1

54.2

-

-

(
Liu and Yuan 2018
)

Heatmap+P

Late

94.5

-

91.7

-

-

-

-

-

(
Zhu et al. 2018
)

RGB+P

Late

92.5

-

94.3

-

-

-

-

-

(
Baradel et al. 2018
)

S+P

Late

-

-

86.6

-

-

-

-

-

(
Luvizon et al. 2018
)

RGB+P

Late

-

-

85.5

-

-

-

-

-

(
Khaire et al. 2018
)

RGB+S+D

Late

95.1

-

-

-

-

-

-

-

(
Shahroudy et al. 2016b
)

RGB+D

Middle

-

-

74.9

-

-

-

97.5

-

(
Shaikh et al. 2022
;
Shaikh et al. 2024
)

RGB+Audio

Middle

-

-

-

-

-

-

-

86.7

Convolutional Neural Network (CNN)-based Multimodal Human Activity Recognition (MHAR) approaches can leverage the potential of automated feature learning, wherein varying modalities inform and interact with each other through concealed features. For a comparison of these techniques on benchmark datasets, please refer to Table
1
. Some significant studies are summarized below.

Islam et al.
(
Islam and Iqbal 2020
)
introduced a Hierarchical Multimodal Attention-based Human Activity Recognition Algorithm (HAMLET), which uses a unique feature encoder and a multi-head self-attention mechanism for each modality to encode spatio-temporal features. HAMLET employs a multimodal self-attention-based fusion architecture, called Multimodal Atention-based Feature
Fusion (MAT), which blends attention-based fused features with the use of sum (extracted unimodal features are summed
after applying multimodal attention) and concatenate (in this approach the attended multimodal features are concatenated) operations. Action classification is done using a fully connected layer that harnesses the computed multimodal features. Contrastingly,
(
Memmesheimer et al. 2020
)
developed an intuitive methodology that blends different modalities using a matrix concatenation operation, transforming signals into an image for classification via a 2D CNN.

Another distinctive approach put forth by
(
Das et al. 2020
)
, combines spatial embeddings of RGB images and 3D poses using an attention mechanism for extracting superior discriminatory spatio-temporal patterns. Moreover,
(
De Boissiere and Noumeir 2020
)
has proposed the late fusion of feature vectors from infrared (IR) and 3D pose modules, classified by a multi-layer perceptron. Notably,
(
Vaezi Joze et al. 2020
)
introduced an innovative method using intermediate fusion among RGB and 3D pose information through a Multimodal Transfer Module (MMTM). Moreover,
(
Luvizon et al. 2018
)
suggests a multi-task framework for concurrent 2D and 3D estimation from still images and HAR from video sequences.

The work proposed by
(
Perez-Rua et al. 2019
)
used a network architecture search-based approach that applies a progressive algorithm for multi-fusion architecture search and the introduction of new fusion layers. This technique executes an explicit fusion of video and pose modalities via a search algorithm. Alternatively,
(
Das et al. 2019
)
introduced an approach that utilizes a separable spatio-temporal attention model and a pose-driven attention model, late fusing scores by averaging softmax scores.

The Evolution of Pose Estimation Maps (PEM), proposed by
(
Liu and Yuan 2018
)
, employs spatial rank pooling to aggregate the evolution of heatmaps as a body shape evolution image, and body-guided sampling for aggregating the evolution of poses as a body pose evolution image. Complementary features of both images are then probed through CNNs for action classification. Further,
(
Baradel et al. 2018
)
has proposed a method that uses unstructured collections of spatio-temporal glimpses with distributed recurrent tracking. Additionally,
(
Shahroudy et al. 2016b
)
has presented a deep multimodal feature analysis-based learning machine that applies mixed norms for component regularization and group selection for superior classification performance. Lastly,
(
Khaire et al. 2018
)
has used a 5-CNN-streams approach, based on Motion History Image (MHI), Front Depth Motion Maps (DMMs), Side DMM, Top DMM, and Skeleton images, fused at the decision level for action classification.

These various approaches offer different methods of fusing multimodal information and capitalizing on their synergies for improved performance in Human Activity Recognition. They demonstrate versatility and innovation in this rapidly advancing field.

2.3.
Learning Methods in Transformers

We first introduce Transformers here and then illustrate the major trends seen across all Transformer-based architectures that have also been adopted for action recognition.

2.3.1.
Transformers

Figure 3.
The Transformer, as originally proposed in
(
Vaswani et al. 2017
)
, depicted through visualization
(
Selva et al. 2022
)
.

Transformers, as introduced by Vaswani et al.
(
Vaswani et al. 2017
)
, have redefined neural network architectures, primarily due to their superior capability in capturing long-term sequential data dependencies. The Transformer was first proposed as a remedy to some limitations of sequence modeling architectures, originally designed to deal with whole sequences at once (see Fig.
3
), allowing parallelization of some operations (as opposed to RNNs, which are sequential in nature), and reducing the locality bias of traditional networks (such as CNNs). The key innovation in transformers is the self-attention mechanism, which allows the network to selectively focus on different parts of the input sequence based on their relevance to the current task. Different from RNNs, which operate sequentially, and traditional networks like CNNs that have a locality bias, Transformers treat entire sequences in parallel, a major breakthrough enabled by the self-attention mechanism.

The essence of the self-attention mechanism is computing three vectors for every input element: query, key, and value. The mechanism assigns weights to each input element based on its query vector’s similarity, then derives an output from a weighted sum of the value vectors. Additionally, Transformers have other modules such as:
Multi-head attention
: It focuses on multiple sequence parts simultaneously by generating multiple sets of query, key, and value vectors. The outputs are then amalgamated using a linear operation.

(7)

Attention
​
(
Q
,
K
,
V
)
\displaystyle\text{Attention}(Q,K,V)

=
softmax
​
(
Q
​
K
T
d
k
)
​
V
\displaystyle=\text{softmax}\left(\frac{QK^{T}}{\sqrt{d_{k}}}\right)V

(8)

MultiHead
​
(
Q
,
K
,
V
)
\displaystyle\text{MultiHead}(Q,K,V)

=
Concat
​
(
h
​
e
​
a
​
d
1
,
…
,
h
​
e
​
a
​
d
h
)
​
W
O
\displaystyle=\text{Concat}(head_{1},\ldots,head_{h})W^{O}

where

(9)

h
​
e
​
a
​
d
i
=
Attention
​
(
Q
​
W
i
Q
,
K
​
W
i
K
,
V
​
W
i
V
)
head_{i}=\text{Attention}(QW_{i}^{Q},KW_{i}^{K},VW_{i}^{V})

and
W
i
Q
W_{i}^{Q}
,
W
i
K
W_{i}^{K}
,
W
i
V
W_{i}^{V}
, and
W
O
W^{O}
are the learnable weight matrices.
Position-wise feedforward network
: This module applies individual feedforward networks to every sequence element, which enhances the model’s sequential data handling.

(10)

Position-wise FeedForward
​
(
x
)
=
max
​
(
0
,
x
​
W
1
+
b
1
)
​
W
2
+
b
2
\text{Position-wise FeedForward}(x)=\text{max}(0,xW_{1}+b_{1})W_{2}+b_{2}

where
x
x
,
W
1
W_{1}
,
W
2
W_{2}
,
b
1
b_{1}
, and
b
2
b_{2}
are input and learnable parameters.

2.3.2.
Transformers-based Learning Methods

Transformer-based methods have also been explored for multimodal action recognition and have shown promising results. The self-attention mechanism of the Transformer allows it to selectively attend to different modalities and capture their temporal and spatial relationships, leading to improved performance compared to traditional methods
(
Gabeur et al. 2020
)
. Additionally, the use of cross-modal attention can help the model fuse information from different modalities and improve its overall accuracy
(
Zhuang et al. 2017
)
.
Table.
2
presents summary on architectures of some significant transformer-based works in MHAR and refers to various features like Normalization (Norm.), Activation (Act.), Encoder configuration (Heads/Layers), Efficiency (Aggregation, Restriction, and Weight-sharing), Positional-Encoding (PE, with type and strategy), and Backbone networks.

Initial methods such as VATNet
(
Girdhar et al. 2019
)
, CBT
(
Wang et al. 2019b
)
, and ELR
(
Purwanto et al. 2019
)
set a precedent. Their successes paved the way for a surge in transformer-based approaches, revealing a shift from traditional CNN methods, which primarily hinged on spatial features. Most methods either use Pre or Post normalization. When paired with activation, GeLU and ReLU seem to be dominant choices. The varied combinations indicate exploration to pinpoint the most efficient configuration.

The hierarchy of encoder architecture varies across different methods. While some like Actor-T
(
Gavrilyuk et al. 2020
)
showcase simpler design configurations, others like Video Swin
(
Liu et al. 2022
)
take a deep dive with relatively complex design choices. This reflects the diverse depth of transformer models tailored for action recognition. However, notable methods, such as TimeSformer
(
Bertasius et al. 2021
)
, ViViT
(
Arnab et al. 2021
)
, VATT
(
Akbari et al. 2021
)
, MViT
(
Fan et al. 2021
)
, SCT
(
Zha et al. 2021
)
, and others, have indicated the inclusion of aggregation, restriction, and weight-sharing strategies for achieving computational efficiency in design.

A majority of the methods, including ViViT
(
Arnab et al. 2021
)
, SCT
(
Zha et al. 2021
)
, and VTN
(
Wu et al. 2021
)
, adopt positional encoding mechanisms, with strategies ranging from Learnt (L) to Fixed (F) and from Absolute (A) to Relative (R). This underscores the essential role of temporal awareness in videos. For backbone choices, transformer-based methods for MHAR leverage diverse CNN-based backbones to capture spatiotemporal features essential for video understanding. In particular, I3D
(
Carreira and Zisserman 2017
)
is commonly adopted, known for its efficacy in video-related tasks, as seen in models like VATNet
(
Girdhar et al. 2019
)
, LapFormer
(
Kondo 2021
)
, and Actor-T
(
Gavrilyuk et al. 2020
)
. S3D
(
Xie et al. 2017
)
in CBT
(
Wang et al. 2019b
)
efficiently separates spatial and temporal information, enhancing video content understanding. ResNet-50, utilized in LapFormer
(
Kondo 2021
)
, LTT
(
Kalfaoglu et al. 2020
)
, PEMM
(
Lee et al. 2020
)
, TRX
(
Perrett et al. 2021
)
, and GroupFormer
(
Li et al. 2021
)
, provides a strong foundation with residual connections, addressing the vanishing gradient problem. R(2+1)D
(
Tran et al. 2018
)
, found in LTT
(
Kalfaoglu et al. 2020
)
, STiCA
(
Patrick et al. 2021
)
, and GroupFormer
(
Li et al. 2021
)
, decomposes 3D convolutions for efficient spatiotemporal processing. SlowFast
(
Feichtenhofer et al. 2019
)
, employed in PEMM
(
Lee et al. 2020
)
and CATE
(
Sun et al. 2021
)
, captures both slow and fast motion features, suitable for varied temporal dynamics. The use of HRNet
(
Wang et al. 2020a
)
in Actor-T
(
Gavrilyuk et al. 2020
)
emphasizes maintaining high-resolution representations for detailed visual information. Additionally, various methods, including TimeSformer
(
Bertasius et al. 2021
)
, ViViT
(
Arnab et al. 2021
)
, FAST
(
Yu et al. 2021
)
, VATT
(
Akbari et al. 2021
)
, MViT
(
Fan et al. 2021
)
, SCT
(
Zha et al. 2021
)
, CATE
(
Sun et al. 2021
)
, Video Swin
(
Liu et al. 2022
)
, and VTN
(
Wu et al. 2021
)
, employ custom linear layers, showcasing a flexible approach to multimodal feature fusion within the Transformer architecture. Overall, the selection of CNN-based backbones reflects a strategic choice to enhance the ability of the model to recognize human actions in diverse multimodal contexts.

Transformers, despite their successes, struggle with challenges. A scarcity of large-scale multimodal datasets reduces their performance, making valid comparisons with CNN-based methods difficult. Their intricate architecture necessitates high computational power and memory. This underlines an acute need for streamlined transformer architectures.
The success of transformer-based models, driven by their flexibility and depth, has steered the action recognition domain into an experimental phase. Their departures and incorporations from CNN techniques not only notice their distinctive design philosophy but also underline the evolutionary trajectory of video understanding. However, addressing their inherent challenges remains pivotal for yielding their full potential in real-world scenarios.

Table 2.
Summary of significant Transformer-based architectures in MHAR. Abbreviations: Norm: Normalization, Act: Activation (Pre or Post), Encoder: number of Heads/Layers (“-” subsequent blocks and “*” repeated modules), Efficient: “X” for indicating Aggregation,
Restriction and Weight-sharing, PE: Positional-Encoding (L: Learnt/F: Fixed, A:Absolute/R: Relative), T: Tokenization (Patch, Clip, and Frame levels), Backbone (ResNet and
DenseNet is abbreviated). Modality: (Visual, Audio, or Textual), SS: Self-Supervision,
NA: not available, X/“–” applies or not.

Model

Norm.

Act.

Encoder

(H/L)

Backbone

PE

T

Modal.

SS

VATNet
(
Girdhar et al. 2019
)

Post

ReLU

2/3

I3D
(
Carreira and Zisserman 2017
)

Faster R-CNN
(
Ren et al. 2015
)

FA

P

X – –

–

CBT
(
Wang et al. 2019b
)

Post

GeLU

4-1/1-NA

S3D
(
Xie et al. 2017
)

–

C

X – X

X

ELR
(
Purwanto et al. 2019
)

Post

ReLU

8/1

I3D
(
Carreira and Zisserman 2017
)

–

P

X – –

–

LapFormer
(
Kondo 2021
)

Post

ReLU

4/3

RN-50
(
He et al. 2016
)

FA

P

X – –

–

LTT
(
Kalfaoglu et al. 2020
)

Pre

GeLU

8/1

R(2+1)D
(
Tran et al. 2018
)

LA

F

X – –

–

Actor-T
(
Gavrilyuk et al. 2020
)

Post

ReLU

1/1

I3D
(
Carreira and Zisserman 2017
)

HRNet
(
Wang et al. 2020a
)

FA

P

X – –

–

TimeSformer
(
Bertasius et al. 2021
)

Pre

GeLU

12/12

Linear Layer

LA

P

X - -

-

PEMM
(
Lee et al. 2020
)

Post

GeLU

6/12*3

SlowFast
(
Feichtenhofer et al. 2019
)

RN-50
(
He et al. 2016
)

LA

C

X X –

X

ViViT
(
Arnab et al. 2021
)

Pre

GeLU

16/24

Linear Layer

LA

P

X – –

–

FAST
(
Yu et al. 2021
)

Pre

GeLU

NA/8

2 Linear Layers

LA

P

X – –

–

VATT
(
Akbari et al. 2021
)

Pre

GeLU

12/12

Linear Layer

LA

P

X X X

X

MViT
(
Fan et al. 2021
)

Pre

GeLU

1-2-4-8/

1-2-11-2

Linear Layer

LA

P

X – –

–

SCT
(
Zha et al. 2021
)

Pre

GeLU

4-8-8-8/

6-1-1-4

Linear Layer

LA

P

X – –

–

CATE
(
Sun et al. 2021
)

NA

NA

12/4

SlowFast
(
Feichtenhofer et al. 2019
)

–

C

X – –

X

TRX
(
Perrett et al. 2021
)

Pre

NA

1/1

RN-50
(
He et al. 2016
)

FA

F

X – –

–

STiCA
(
Patrick et al. 2021
)

Post

ReLU

4/2

R(2+1)D-18
(
Tran et al. 2018
)

RN-9
(
He et al. 2016
)

LA

C

X X –

X

GroupFormer
(
Li et al. 2021
)

Post

ReLU

8*4

I3D
(
Carreira and Zisserman 2017
)

LA

F

X – –

–

Video Swin
(
Liu et al. 2022
)

Pre

GeLU

3-6-12-24/

2-2-6-2

Linear Layer

LR

P

X – –

–

MAiVAR-T
(
Shaikh et al. 2023
)

Pre

ReLU

3-6-12-24/

2-2-6-2

TSN
(
Wang et al. 2016
)

-

P

X X –

–

VTN
(
Wu et al. 2021
)

Pre

GeLU

12-12/12-3

Linear Layer

LA

P

X – –

–

3.
Datasets

A multitude of datasets have been created to develop and evaluate multimodal human activity recognition (MHAR) methods. In Table
3
, a series of benchmark datasets are listed along with their attributes such as scale, number of categories and provided modalities. For RGB-based MHAR, UCF101
(
Soomro et al. 2012
)
, HMDB51
(
Kuehne et al. 2011
)
, and Kinectis-400
(
Kay et al. 2017
)
are widely used as benchmark datasets, along with Kinetics-600
(
Carreira et al. 2018
)
, Kinetics-700
(
Carreira et al. 2019
)
, EPIC-KITCHENS-55
(
Damen et al. 2018
)
, and ActivityNet
(
Fabian Caba Heilbron et al. 2015
)
. Additionally, the datasets MSR-DailyActivity3D
(
Wang et al. 2014b
)
and Northwestern-UCLA
(
Wang et al. 2014a
)
are widely used for depth-based MHAR. Moreover, for MHAR, large benchmark datasets suitable for fusion and co-learning of different modalities include NTU RGB+D
(
Shahroudy et al. 2016a
)
, NTU RGB+D 120
(
Liu et al. 2019
)
, and MMAct
(
Kong et al. 2019
)
. The datasets UTD-MHAD
(
Chen et al. 2015
)
, and PKU-MMD
(
Liu et al. 2017
)
are also popularly used.

MHAR datasets are generally created using different sensor perspectives and settings. For instance, there are single-view action datasets, multi-view action datasets and multi-person/interaction action datasets. In the single-view action datasets, a specific single viewpoint is captured. In the multi-view datasets, two or more viewpoints of a single action are captured in the frame sequences. In multi-person/interaction action datasets, action/activity is normally performed between two or more people
(
Zhang et al. 2016
)
. Additionally, action recognition datasets are normally available in different modalities, namely RGB, depth videos, skeleton joint positions, inertial sensor signals, and IR. These modalities often allow capturing different aspects of information for the same action scenario. As can be seen in Table
3
, the datasets also vary considerably in terms of the size and number of classes. As there is no public platform similar to YouTube available for MHAR dataset generation, MHAR currently lacks large-scale benchmark datasets when compared to video-based action recognition.

Table 3.
Public multimodal datasets. Notations : Seg: Segmented, Con: Continuous, D: Depth, S: Skeleton, Au: Audio, Ac: Accelerometer, IR: Infrared, #: Number of.
* action classes with
>
>
50 samples
.
Note: Some RGB datasets are considered as multimodal because different modalities are from it such as audio, skeleton, and optical flow as well as some representations from compressed video formats.

Dataset

Year

Data Source

Type

Modality

#Actions

#Subjects

#Samples

#Views

CMU MoCap
(
Hodgins 2021
)

'01

MoCap

Seg

RGB,S

45

144

2,235

1

MSR-Action3D
(
Wang et al. 2014b
)

'10

Kinect v1

Seg

S,D

20

10

567

1

RGBD-HuDaAct
(
Ni et al. 2011
)

'11

Kinect v1

Seg

RGB+D

13

30

1,189

1

HMDB-51
(
Kuehne et al. 2011
)

'11

YouTube

Seg

RGB

51

-

6,766

1

MSR-DailyActivity3D
(
Wang et al. 2014b
)

'12

Kinect v1

Seg

RGB,D,S

16

10

320

1

UCF-101
(
Soomro et al. 2012
)

'12

YouTube

Seg

RGB

101

-

13,320

1

UT-Kinect
(
Xia et al. 2012
)

'12

Kinect v1

Seg

RGB,D,S

10

10

200

1

SBU Kinect
(
Yun et al. 2012
)

'12

Kinect v1

Seg

RGB,D,S

7

8

300

1

3D Action Pairs
(
Oreifej and Liu 2013
)

'13

Kinect v1

Seg

RGB,D,S

12

10

360

1

Berkeley MHAD
(
Ofli et al. 2013
)

'13

MoCap + Kinect v1

Seg

RGB,D,S,Au,Ac

12

12

660

4

NW-UCLA
(
Wang et al. 2014a
)

'14

Kinect v1

Seg

RGB,D,S

10

10

1,475

3

UTD-MHAD
(
Chen et al. 2015
)

'15

Kinect v1 + Inertial

Seg

RGB,D,S

27

8

861

1

NTU RGB+D 60
(
Shahroudy et al. 2016a
)

'16

Kinect v2

Seg

RGB,D,S,IR

60

40

56,880

80

PKU-MMD
(
Liu et al. 2017
)

'17

Kinect v1

Con

RGB,D,S,IR

51

66

1,076

3

ActivityNet-200
(
Fabian Caba Heilbron et al. 2015
)

'16

YouTube

Con

RGB,X

200

-

19,994

1

Kinetics 400
(
Kay et al. 2017
)

'17

YouTube

Seg

RGB

400

-

306,245

1

Kinetics 600
(
Carreira et al. 2018
)

'18

YouTube

Seg

RGB

600

-

495,547

1

AVE
(
Tian et al. 2018
)

'18

YouTube

Seg

RGB,Audio

28

-

4,143

1

EPIC Kitchen 55
(
Damen et al. 2018
)

'18

GoPro Hero 5

Con

RGB

149*

32

39,596

1

NTU RGB+D 120
(
Liu et al. 2019
)

'19

Kinect v2

Seg

RGB,D,S,IR

120

106

114,480

155

Toyota-SH
(
Das et al. 2019
)

'19

Kinect v1

Seg

RGB,D,S

31

18

16,115

1

Kinetics 700
(
Carreira et al. 2019
)

'19

YouTube

Seg

RGB

700

-

650,317

1

EPIC Kitchen 100
(
Damen et al. 2022
)

'20

GoPro Hero 7

Con

RGB,Flow

4,053

37

89,977

1

MHAR datasets contain continuous (Con) or segmented (Seg) videos as reported in Table
3
. The continuous video datasets comprise more than one action in a single video or frame sequence. They are mainly used for localization, detection, and prediction of actions. The techniques covered in this article generally focus on segmented datasets, which have complete actions in individual data instances. These datasets can be easily used for the evaluation of action recognition algorithms. Below, we discuss different aspects of commonly used segmented datasets.

The CMU MoCap
(
Hodgins 2021
)
is a graphics motion-capture database. It is one of the earliest sources of action data, that covers a variety of actions, including the interaction between two subjects, locomotion with uneven terrains, sports and several other human actions. CMU MoCap is a segmented dataset with 45 classes of action performed on 144 subjects containing 2,235 video samples. This dataset has a single line of sight, and it contains RGB and skeleton modalities.
The MSR-Action3D
(
Wang et al. 2014b
)
is the first RGB+D action dataset captured by the Kinect sensors. It contains 20 actions performed by ten subjects. This dataset requires some post-processing to remove the background. In particular, the right arms and legs are known to be preferred in this dataset for performing several related actions.

The MSR-DailyActivity3D dataset
(
Wang et al. 2014b
)
is collected by Microsoft and Northwestern University, focusing on daily activities in a living room. A total of 16 actions are performed by ten subjects while sitting on a sofa or standing close to a sofa. This dataset functions as a good baseline for multimodal video analysis, providing a small number of classes. The dataset provides RGB, depth and skeleton data modalities.

The Berkeley Multimodal Human Action Database (Berkeley MHAD)
(
Ofli et al. 2013
)
is the first dataset to be captured with five different modalities. One optical MoCap system, four multi-view cameras, two Kinect v1 cameras, six wireless accelerometers, and four microphones are used in the creation of this dataset. Berkeley MHAD contains eleven actions performed by seven male and five female subjects, each varying in terms of style and speed. These actions are divided into the categories of: (1) full body movement actions, e.g., jumping jacks, throwing; (2) high dynamics in the upper extremity actions, e.g., waving hands, clapping hands; and (3) high dynamics in the lower extremity actions, e.g., sitting down, standing up. Audio and accelerometer data are two rare modalities in this dataset.

The NTU RGB+D 120
(
Liu et al. 2019
)
dataset, as highlighted in Table
3
, is currently one of the largest action recognition datasets in terms of the number of samples per action. It has RGB, depth, skeleton, and IR modalities, captured with Kinect sensor v2. The dataset has more than 114,000 action video samples and up to 8 million image frames. The dataset contains 120 classes of daily and health-related actions performed by 106 subjects aged around 10 to 57 years, with different cultural backgrounds. Further, it consists of 155 different camera viewpoints. This dataset is a strong candidate for evaluating the scalability and multimodal nature of algorithms.

The Toyota-Smarthome (Toyota-SH)
(
Das et al. 2019
)
dataset contains approximately 16,100 unscripted video samples with 31 action classes performed by 16 subjects. This dataset is one of the most recent action recognition datasets, with three different scenes and seven camera viewpoints. Toyota-SH dataset provides several real-world challenges such as improvisational acting, camera framing, composite activities, multi-view and the same activity using different objects. A unique feature of this dataset is that its actions are performed by subjects who did not receive any information regarding how to perform them. This dataset provides RGB, depth and skeleton modalities.

4.
Fusion Methods

To this end, recent
works have started to explore different fusing methods to incorporate multi-sensory data streams effectively. Compared with
typical CNNs, a Transformer is naturally appropriate for
multi-stream data fusion because of its nonspecific embedding
and dynamically interactive attention mechanisms. Multimodal data fusion refers to the mixing of features from data of various modalities. The objective of multimodal data fusion is to achieve better accuracy than a single modality. Data fusion supports diversity, enhancing the uses, advantages and analysis of ways that cannot be achieved through a single modality. Merging different modalities, as in
(
Yin et al. 2022
)
, have many benefits such as enhanced signal-to-noise ratio, improved confidence, increased robustness and reliability, enhanced resolution, better precision and discrimination as well as robustness against interference
(
Aguileta et al. 2019
)
.

In the context of MHAR, the choice of formulation depends on data characteristics and specific task requirements, necessitating experimentation and empirical evaluation to identify the most effective fusion strategy for a given multimodal dataset. Researchers explore diverse architectures and fusion methods to optimize accuracy and robustness in MHAR systems. Fusion techniques can be deployed at different stages in the action recognition process to acquire combinations of distinct information. These fusion approaches depend on the type of data and technique used for action recognition. Multimodal data fusion approaches could be classified into classical-machine-learning-based and deep learning-based, where latter can be further divided into CNNs- and Transformers-based, according to architectural choices. Various fusion approaches are discussed in the following subsections.

4.1.
Fusion in CNNs

Data fusion in deep-learning-based MHAR techniques can be applied at similar stages as in classical-machine-learning-based approaches. In the context of fusion in MHAR, a common strategy involves employing modality-specific networks, where separate CNNs process each modality independently, and their outputs are combined through fully connected layers or other fusion techniques. Additionally, for video-based action recognition, 3D CNNs capture spatial and temporal features simultaneously, offering a natural approach to fuse information across space and time for multiple modalities.

For video-based multimodal action recognition, the utilization of 3D CNNs, which capture spatiotemporal information simultaneously, presents an effective approach. This mirrors the approach often employed in single-modality video action recognition
(
Korbar et al. 2019
)
and provides a natural means of fusing information over both space and time.

To accommodate heterogeneous modalities, studies have adopted the practice of utilizing separate CNNs for each modality
(
Wang et al. 2012
)
. Subsequently, the outputs from these modality-specific networks are combined using fusion layers, often comprising fully connected layers or other fusion techniques.

the incorporation of temporal convolution layers
(
Carreira and Zisserman 2017
)
directly models the temporal dynamics within the data. This is critical for recognizing the temporal evolution of actions, a key factor in accurate action recognition.

However, depending on the nature of the deep neural network, data-level, feature-level and decision-level, techniques are often referred to as “early”
(
Feichtenhofer et al. 2016
)
, “middle”
(
Roitberg et al. 2019
)
or “intermediate”
(
Patel et al. 2018
)
, and “late” fusion. Table
1
presents some common deep-learning-based MHAR methods that use different data fusion approaches, which are discussed in the following subsections.

4.1.1.
Early Fusion

The early-fusion approach captures information and combines it at the feature level
(
Feichtenhofer et al. 2016
)
. In this case, features from different modality sources are concatenated in initial layers into an aggregated feature, which will then be used by the later layers for classification. As this aggregated feature consists of many features, it increases the training and classification time. However, these large aggregates and suitable learning techniques can offer much better recognition performance in the end. In Fig.
4
, the initial process of feature fusion is illustrated as an example of early fusion.

4.1.2.
Middle Fusion

The middle-fusion approach merges the features extracted from raw data in the entire neural network so that the higher layers have access to more global information
(
Roitberg et al. 2019
;
Patel et al. 2018
)
. A CNN-based operation is performed to compute the weights and extend the connectivity of all layers. Middle-level fusion attempts to exploit the advantages of early and late fusion in a common framework.

4.1.3.
Late Fusion

The late-fusion approach indicates a combination of the action information at the deepest layers in the network i.e., after the classification. For example, a typical MHAR network architecture consists of two separate CNN-based networks with shared parameters up to the last convolution layer. The outputs of the last convolution layer of these two separate network streams are processed to the fully connected layer. This step predicts the final output after fusing the classification scores and considers individual class labels at the score layer
(
Shaikh et al. 2022
)
. Different decision rules are deployed to fuse the scores at this stage. Although late fusion ignores some interactions between modality, it adds more simplicity and flexibility in making final decisions when one or more modalities is missing. The late fusion approach has been relatively successful in most MHAR architectures. As evident in Fig.
4
, the stage of Probability fusion is an example of late fusion in deep-learning-based methods.

Figure 4.
An example of multi-level fusion in deep-learning-based action recognition, where the red dotted lines represent four integration points corresponding to different multimodal fusion methods examined. For a specific
integration point, the network is duplicated for
K
K
different modalities, concatenate the features at the integration point, and the network
after the integration point remain unchanged. (adapted from
(
Long et al. 2018
)
)

4.2.
Fusion in Transformers

Transformer-based techniques are commonly used for multimodal data fusion. On the one hand, the fusion process typically involves concatenating input sequences from all modalities, or employing a form of cross-attention mechanism (as shown in Fig.
5c
). Attention mechanisms dynamically weigh the importance of features from various modalities, allowing the network to focus on task-relevant information and effectively fuse data. The fusion process can occur at different stages of the architecture, including early, middle, or late stages. In this paper, we discuss both the how and where aspects of the fusion process in Transformer-based multimodal data fusion techniques.

4.2.1.
Early Fusion

This subsection discusses Early Fusion techniques which involve methods in which the fusion occurs prior to input being supplied to the encoder.

Encoder Fusion (EF)

Prior to being fed into the encoder, token embeddings from different modalities are concatenated either in a sequence-wise manner as in
(
Li et al. 2020
;
Sun et al. 2019b
;
Liu et al. 2021
)
(refer to Fig.
5a
) or in a channel-wise manner as in
(
Fang et al. 2019
)
. The former can be likened to the approach of BERT
(
Devlin et al. 2018
)
in dealing with pairs of language sentences. To differentiate the tokens belonging to each modality,
(
Sun et al. 2019b
)
separator tokens are used, akin to the [SEP] token in BERT, which originally indicates the start of a new sentence but here signals that the succeeding tokens are from another modality. However, most video transformers (VTs) indicate the modality of a given token by summing or concatenating learned modality embeddings in a manner similar to the addition of position embeddings. Encoder fusion significantly increases the computational cost of the self-attention operation up to
O
⁡
(
(
T
1
+
…
+
T
m
)
2
)
,
O((T_{1}+...+T_{m})^{2}),
where
M
M
is the number of modalities and
T
m
T_{m}
is the number of tokens in the
m
m
-th modality. It is also worth noting here that we not only explore how but also where the fusion is integrated into the architecture, namely at early, middle, or late stages.

(a)

(b)

(c)

(d)

Figure 5.
Four main types of performing multimodal fusion in Transformers (adapted from
(
Selva et al. 2022
)
).

Hierarchical Encoder Fusion (HEF)

In multimodal learning, Encoder fusion can be performed hierarchically by initially enhancing modality-specific token embeddings on individual encoders, concatenating their outputs, and sending them to a multimodal encoder
(
Lee et al. 2020
;
Sun et al. 2019a
;
Pashevich et al. 2021
)
(see Fig.
5b
). This approach enables intra-modal information to be dealt with before modelling the inter-modal patterns, which is advantageous when handling modalities that are not closely related or accurately aligned at the input level. Experimentation is necessary to determine when to do this. Furthermore, the computational cost is lower than that of the previous encoder fusion by a constant factor, depending on the number of layers in both the unimodal and the multimodal encoders. Although using multiple encoders increases the number of parameters, this can be mitigated by weight sharing
(
Lee et al. 2020
)
.

4.2.2.
Middle Fusion

This subsection discusses Middle Fusion techniques which involve methods in
which the fusion occurs before combining the
outputs of the different encoders.

Cross-Attention Fusion (CAF)

In the process of modalities fusion using CA, one modality seeks information from another modality to provide context (refer to Fig.
5c
). This simple idea and its flexibility have led to different uses in various studies. In
(
Kim et al. 2018
;
Iashin and Rahtu 2020b
)
, separate encoder-decoders are proposed for each modality, and each cross-attends to text embeddings before combining the outputs of the different encoders. The study of
(
Jaegle et al. 2021
)
proposes only one stream that keeps augmenting a small set of latent embeddings by repeatedly cross-attending to the same very long multimodal input sequence of minimal embeddings. In this case, CA layers are interleaved with SA layers that refine the cross-attended information.
(
Zhu and Yang 2020
)
proposes a three-stream Transformer, where the central one cross-attends to the other two, and these, in turn, attend to the embeddings generated by their respective opposite previous cross-attentions. In contrast,
(
Camgoz et al. 2020
)
also uses three streams, one per modality. However, the fusion is achieved within a master stream that substitutes its SA by asymmetric cross-attention over the other two simultaneously (concatenating both sets of keys and values).

Co-Attention Fusion (CoAF)

In contrast to the cross-attention fusion, co-attention involves parallel augmentation of the two modalities by attending to each other’s embeddings. In this approach, the CA sub-layer in each stream generates the queries from its own embeddings, whereas keys and values are obtained from the other stream (see Fig.
5d
). The ViLBERT model
(
Lu et al. 2019
)
originally proposed this method for images and language, where it has since been adopted by several video-related works
(
Iashin and Rahtu 2020a
;
Su et al. 2021
;
Curto et al. 2021
)
. While some works entirely replaced SA with CA
(
Su et al. 2021
;
Curto et al. 2021
)
, others retained it
(
Iashin and Rahtu 2020a
)
. In contrast to encoder fusion, co-attention reduces the computational cost to
O
⁡
(
(
T
1
​
T
2
)
2
)
O((T_{1}T_{2})^{2})
. Additionally,
(
Zhang et al. 2021
)
has advocated for co-attending to each other and self-attending to themselves as a means of maintaining intra-modal and inter-modal dynamics separately.

4.2.3.
Late Fusion

All the aforementioned strategies for fusing modalities offer various early and middle fusion options. However, an alternative approach is to perform a late fusion of modalities. This involves running different modalities through parallel encoders and combining their outputs. In classification problems, these outputs could be class score distributions
(
Gavrilyuk et al. 2020
)
, as is typically done for Two-stream ConvNets
(
Simonyan and Zisserman 2014
)
. Although this strategy may be suboptimal for Transformers, it may still be useful when there is limited training data available. However, late fusion could be considered by concatenating the augmented aggregation token as proposed by
(
Kim et al. 2018
)
.

There are two concatenation-based approaches, namely Encoder Fusion (EF) and Hierarchical Encoder Fusion (HEF), the former of which concatenates the modalities at an early stage while the latter concatenates them at a middle stage, which is more computationally intensive but eliminates the need to experimentally determine the fusion point. In contrast, CAF and CoAF are based on cross-attention mechanisms and are inherently limited to fusing only two modalities. Nonetheless, recent studies have attempted to extend these methods to more than two modalities, as evidenced by the cumbersome approach proposed by
(
Zhu and Yang 2020
)
in their ActBERT model. Compared to EF and HEF, both CAF and CoAF are more efficient, but they still suffer from the issue of identifying the optimal fusion point. However, middle fusion may be useful when the input modalities are asynchronous or differ in nature.

5.
Challenges, Future Directions And Limitations

In this section of the survey article, we will delve into the potential challenges, future directions associated with the field under discussion and limitations of this survey. By exploring the challenges, we aim to shed light on the hurdles that must be overcome to make further progress. Further, by discussing future directions, we aim to provide insights into the exciting new avenues for research and development that lie ahead. Finally, by acknowledging the limitations of this survey, we aim to provide a balanced and realistic view of its current scope and limitations.

5.1.
Challenges and Future Directions

Future of Datasets:
Extensive and comprehensive datasets hold significant importance for the advancement of MHAR, particularly for MHAR methods based on deep learning. Several factors, such as the magnitude, diversity, applicability, and modality type, indicate its quality. Despite the availability of numerous datasets that have significantly progressed the MHAR field, the development of novel benchmark datasets is still necessary to further promote research in this domain. For instance, most existing multimodal datasets have been obtained in controlled settings with voluntary participants performing actions. Consequently, gathering multimodal data from uncontrolled environments to establish substantial and challenging benchmarks for enhancing MHAR’s practical applications would be beneficial. Moreover, constructing large and complex datasets for recognizing each person’s actions in crowded settings warrants further exploration.

Multimodal Learning:
Earlier discussions have outlined various multimodal learning methods, including multimodal fusion and cross-modality transfer learning, proposed for MHAR. Multimodal data fusion, which can complement each other, has been shown to enhance MHAR performance while co-learning can address data deficiencies in certain modalities. Nevertheless, as
(
Wang et al. 2020b
)
has highlighted, several challenges impede the effectiveness of many existing multimodal methods, such as the risk of overfitting. This underscores the need to develop more effective strategies for multimodal fusion and co-learning in MHAR. This challenge requires further exploration and innovation to develop advanced techniques for addressing the limitations of the existing state-of-the-art multimodal methods.

Labelled Data Scarcity in Unsupervised and Semi-Supervised Learning:
The application of supervised learning methods, particularly deep learning-based ones, typically necessitates a substantial amount of labelled data for model training, which can be expensive to obtain. Conversely, unsupervised and semi-supervised learning techniques
(
Alayrac et al. 2016
;
Singh et al. 2021
;
Song et al. 2021
)
have the potential to utilize unlabelled data to train models, thereby mitigating the requirement for extensive labelled datasets. Since unlabelled action samples are often more easily collected than labelled ones, unsupervised and semi-supervised MHAR has emerged as a crucial research direction that merits further investigation.

Generalization:
It is acknowledged that while some works have investigated the generalization capabilities of Transformers, this issue remains open and understudied, particularly in the case of VTs. Although some studies have examined the generalization of Transformers in natural language processing, such as out-of-distribution (OOD) generalization
(
Hendrycks et al. 2020
)
and cross-modal transfer learning with minimal fine-tuning
(
Lu et al. 2021
)
, the analysis of the generalization capabilities of VTs is scarce. To the best of our knowledge, only one work has thoroughly investigated the topic of generalization in VTs
(
Zhang et al. 2022
)
.

Self-Supervised Approaches:
There is a need for further investigation about video self-supervised tasks and their applicability to VTs. While contrastive techniques dominate self-supervised approaches, there are several untapped possibilities for utilizing self-supervised video tasks on VTs, including temporal consistency
(
Fernando et al. 2017
)
, interframe predictability
(
Han et al. 2019
)
, geometric transformations
(
Kim et al. 2019
)
and motion statistics
(
Wang et al. 2019a
)
. Additionally, the recent emergence of self-supervised vision tasks like SimCLR
(
Chen et al. 2020
)
and Barlow Twins
(
Zbontar et al. 2021
)
for images, or BraVE
(
Recasens et al. 2021
)
for video, warrants further exploration for their application to VTs.

Online and Mobile MHAR:
While some studies have explored the use of deep MHAR on mobile devices, such as smartphones and watches, they still face significant limitations in online and mobile deployment
(
Lane et al. 2015
;
Bhattacharya and Lane 2016
)
. In these approaches, the model is trained offline on a remote server, and the mobile device only employs the trained model, which is neither real-time nor efficient for incremental learning. Two potential solutions to address this challenge are to reduce the communication cost between mobile and server and to improve the computing ability of mobile devices.

Collaboration in Deep-Shallow MHAR:
High computational requirements of deep models are often a bottleneck, rendering them unsuitable for wearable devices. Contrastingly, shallow neural networks (NN) and traditional pattern recognition (PR) methods are not able to achieve high-performance levels. Therefore, a middle ground needs to be found by developing lightweight deep models that can still achieve high-performance levels while being efficient enough to run on mobile devices. In this regard, the collaborative efforts of deep and shallow models have the potential to provide accurate and lightweight MHAR solutions. However, several issues need to be addressed, such as how to effectively share parameters between deep and shallow models. Moreover, the existing offline training of deep models hinders real-time execution, necessitating the need for online training approaches that can enhance the adaptability of the model to new environments.

Effective Transfer Learning:
The process of data annotation via transfer learning is executed through the utilization of labelled data from auxiliary domains
(
Wang et al. 2017
)
. Various factors related to human activity can be leveraged as supplementary information through deep transfer learning. This approach poses several challenges, including the sharing of weights between networks, the effective exploitation of knowledge between activity-related domains, and the identification of relevant domains. These issues require resolution in order to facilitate the effective application of transfer learning methodologies.

Hybrid Sensor and Context-Aware MHAR:
The comprehensive information obtained from hybrid sensors is of considerable utility in discerning fine-grained activities
(
Vepakomma et al. 2015
)
. Particular emphasis must be placed on the recognition of such activities through the collaborative utilization of hybrid sensors. In contrast, context refers to any information that can be employed to describe the circumstances surrounding an entity
(
Abowd et al. 1999
)
. Contextual data sources, such as Wi-Fi, Bluetooth, and GPS, can provide valuable insight into environmental factors associated with a given activity. The effective exploitation of contextual information can greatly enhance the ability to recognize both user states and specific activities.

Beyond Activity Recognition:

The recognition of activities frequently represents the initial stage in a variety of applications. For example, professional skill assessment is necessary for fitness training, while smart home assistants are indispensable components of healthcare services. Early efforts in activity recognition include research on climbing assessment
(
Khan et al. 2015
)
. Recent investigations suggest that leveraging the expertise of crowds can substantially facilitate the task
(
Prelec et al. 2017
)
. Crowdsourcing offers a means of annotating unlabelled activities by utilizing the collective abilities of a large group. In addition to passive label acquisition, researchers may also develop more sophisticated and privacy-preserving methodologies to collect valuable labels. As deep learning continues to advance, activity recognition applications can be expanded beyond simple recognition tasks.

5.2.
Limitations

Our study was limited to research papers published within the last 10 years (2012-2022) and only included papers in English, potentially excluding some studies in other languages. Only studies that utilized visual data were considered and others, such as those using olfactory, infrared, or tactile data, were outside the scope of this research. A potential limitation is publication bias, which may lead to an overestimation of the advantages of using multiple forms of data for analysis. The studies reviewed employed different input methods and evaluated various methods for recognizing actions on different datasets, making direct comparison of results difficult. Additionally, not all articles provided statistical confidence intervals, making it challenging to compare their findings.

6.
Conclusion

In this survey, we have conducted a comprehensive analysis and synthesis of the primary advancements and emerging trends in adapting Convolutional Neural Networks (CNNs) and Transformers for multimodal recognition of human actions. Drawing upon the existing literature, we have devised a robust taxonomy for CNN and Transformer architectures based on their modality, intended task and overall structure. Furthermore, we have investigated diverse techniques for embedding, encoding and fusing different multimodal representations to achieve more accurate recognition of human actions.

In addition, we have provided a comparative analysis of the leading approaches on different datasets and suggested key design modifications necessary for improving their performance. Moreover, we have discussed the current trends and potential challenges associated with the different components of the multimodal action recognition pipeline. Despite the considerable progress that has been made, the potential of multimodal fusion for action recognition remains largely unexplored, and there remain several significant challenges that need to be addressed. Finally, we express a keen interest in comparing Transformer architectures to CNNs for multimodal understanding, given the tremendous promise of global-based learning methods.

Acknowledgements.

This work is partially funded by Edith Cowan University (ECU) and the Higher Education Commission (HEC) of Pakistan under Project #PM/HRDI-UESTPs/UETs-I/Phase-1/Batch-VI/2018. Naveed Akhtar is the recipient of an Australian Research Council Discovery Early Career Researcher Award (project number DE230101058) funded by the Australian Government.

References

(1)

Abowd et al
.
(1999)

Gregory D Abowd, Anind K Dey, Peter J Brown, et al
.
1999.

Towards a better understanding of context and context-awareness. In
Int. symposium on handheld and ubiquitous computing
. Springer, 304–307.

Aggarwal and Xia (2014)

J.K. Aggarwal and Lu Xia. 2014.

Human Activity Recognition from 3D Data: A Review.

Pattern Recognition Letters
48 (2014), 70–80.

Aguileta et al
.
(2019)

Antonio A. Aguileta, Ramon F. Brena, Oscar Mayora, Erik Molino-Minero-Re, and Luis A. Trejo. 2019.

Multi-Sensor Fusion for Activity Recognition : a Survey.

Sensors
19 (Sept. 2019), 3808.

Ahmad and Conci (2019)

Kashif Ahmad and Nicola Conci. 2019.

How Deep Features Have Improved Event Recognition in Multimedia: A Survey.

ACM Transactions on Multimedia Computing Communication Applications
15, 2, Article 39 (jun 2019), 27 pages.

Akbari et al
.
(2021)

Hassan Akbari, Liangzhe Yuan, Rui Qian, et al
.
2021.

Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text.

Advances in Neural Information Processing Systems
34 (2021), 24206–24221.

Alayrac et al
.
(2016)

Jean-Baptiste Alayrac, Piotr Bojanowski, Nishant Agrawal, et al
.
2016.

Unsupervised learning from narrated instruction videos. In
Proc. of the IEEE Int. Conf. on Computer Vision and Pattern Recognition
. IEEE, 4575–4583.

Arnab et al
.
(2021)

Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lučić, and Cordelia Schmid. 2021.

ViViT: A video vision transformer. In
Proc. of the Int. Conf. on Computer Vision
. IEEE, 6836–6846.

Atrey et al
.
(2010)

Pradeep K Atrey, M Anwar Hossain, Abdulmotaleb El Saddik, and Mohan S Kankanhalli. 2010.

Multimodal fusion for multimedia analysis: a survey.

Multimedia systems
16 (2010), 345–379.

Baradel et al
.
(2018)

Fabien Baradel, Christian Wolf, Julien Mille, et al
.
2018.

Glimpse Clouds: Human Activity Recognition from Unstructured Feature Points. In
Proc. of the IEEE Int. Conf. on Computer Vision and Pattern Recognition
. IEEE, 469–478.

Batchuluun et al
.
(2019)

Ganbayar Batchuluun, Dat Tien Nguyen, Tuyen Danh Pham, Chanhum Park, and Kang Ryoung Park. 2019.

Action recognition from thermal videos.

IEEE Access
7 (2019), 103893–103917.

Bertasius et al
.
(2021)

Gedas Bertasius, Heng Wang, and Lorenzo Torresani. 2021.

Is space-time attention all you need for video understanding?. In
Proc. of the Int. Conf. on Machine Learning
, Vol. 2. 4.

Bhattacharya and Lane (2016)

Sourav Bhattacharya and Nicholas D Lane. 2016.

From smart to deep: Robust activity recognition on smartwatches using deep learning. In
Proc. of the IEEE Int. Conf. on Pervasive Computing and Communication Workshops
. IEEE, 1–6.

Bruce et al
.
(2020)

XB Bruce, Yan Liu, and Keith CC Chan. 2020.

Vision-Based Daily Routine Recognition for Healthcare with Transfer Learning.

Int. Journal of Biomedical and Biological Engineering
14 (2020), 178–186.

Camgoz et al
.
(2020)

Necati Cihan Camgoz, Oscar Koller, Simon Hadfield, and Richard Bowden. 2020.

Multi-channel transformers for multi-articulatory sign language translation. In
European Conf. on Computer Vision
. Springer, 301–319.

Carfagni et al
.
(2019)

Monica Carfagni, Rocco Furferi, Lapo Governi, et al
.
2019.

Metrological and Critical Characterization of the Intel D415 Stereo Depth Camera.

Sensors
19, 3 (Jan. 2019), 489.

Carreira et al
.
(2018)

Joao Carreira, Eric Noland, Andras Banki-Horvath, Chloe Hillier, and Andrew Zisserman. 2018.

A short note about kinetics-600.

arXiv:1808.01340
(2018).

Carreira et al
.
(2019)

Joao Carreira, Eric Noland, Chloe Hillier, and Andrew Zisserman. 2019.

A short note on the kinetics-700 human action dataset.

arXiv:1907.06987
(2019).

Carreira and Zisserman (2017)

Joao Carreira and Andrew Zisserman. 2017.

Quo Vadis, action recognition? a new model and the kinetics dataset. In
Proc. of the IEEE Int. Conf. on Computer Vision and Pattern Recognition
. 6299–6308.

Chen et al
.
(2015)

C. Chen, R. Jafari, and N. Kehtarnavaz. 2015.

UTD-MHAD: A multimodal dataset for human action recognition utilizing a depth camera and a wearable inertial sensor. In
Proc. of the Int. Conf. on Image Processing
. IEEE, 168–172.

Chen et al
.
(2017)

Chen Chen, Roozbeh Jafari, and Nasser Kehtarnavaz. 2017.

A survey of depth and inertial sensor fusion for human action recognition.

Multimedia Tools Applications
76 (2017), 4405–4425.

Chen et al
.
(2020)

Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020.

A simple framework for contrastive learning of visual representations. In
Proc. of the Int. Conf. on Machine Learning
. PMLR, 1597–1607.

Curto et al
.
(2021)

David Curto, Albert Clapés, Javier Selva, Sorina Smeureanu, Julio Junior, CS Jacques, David Gallardo-Pujol, Georgina Guilera, David Leiva, Thomas B Moeslund, et al
.
2021.

Dyadformer: A multi-modal transformer for long-range modeling of dyadic interactions. In
Proc. of the Int. Conf. on Computer Vision
. IEEE, 2177–2188.

Damen et al
.
(2022)

Dima Damen, Hazel Doughty, Giovanni Maria Farinella, , Antonino Furnari, Jian Ma, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray. 2022.

Rescaling Egocentric Vision.

Int. Journal of Computer Vision
130, 1 (2022), 33–55.

Damen et al
.
(2018)

Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray. 2018.

Scaling Egocentric Vision: the EPIC-KITCHENS Dataset. In
Proc. of the European Conf. on Computer Vision
. Springer.

Das et al
.
(2019)

Srijan Das, Rui Dai, Michal Koperski, Luca Minciullo, Lorenzo Garattoni, Francois Bremond, and Gianpiero Francesca. 2019.

Toyota Smarthome: Real-World Activities of Daily Living. In
Proc. of the IEEE Int. Conf. on Computer Vision
. 833–842.

Das et al
.
(2020)

Srijan Das, Saurav Sharma, Rui Dai, et al
.
2020.

VPN: Learning video-pose embedding for activities of daily living. In
Proc. of the European Conf. on Computer Vision
. Springer, 72–90.

De Boissiere and Noumeir (2020)

A. M. De Boissiere and R. Noumeir. 2020.

Infrared and 3D skeleton feature fusion for RGB-D action recognition.

IEEE Access
8 (2020), 168297–168308.

Devlin et al
.
(2018)

Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018.

BERT: Pre-training of deep bidirectional transformers for language understanding.

arXiv:1810.04805
(2018).

Fabian Caba Heilbron et al
.
(2015)

Bernard Ghanem Fabian Caba Heilbron, Victor Escorcia et al
.
2015.

ActivityNet: A Large-Scale Video Benchmark for Human Activity Understanding. In
Proc. of the IEEE Int. Conf. on Computer Vision and Pattern Recognition
. IEEE, 961–970.

Fan et al
.
(2021)

Haoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li, Zhicheng Yan, Jitendra Malik, and Christoph Feichtenhofer. 2021.

Multiscale vision transformers. In
Proc. of the Int. Conf. on Computer Vision
. IEEE, 6824–6835.

Fang et al
.
(2019)

Kuan Fang, Alexander Toshev, Li Fei-Fei, and Silvio Savarese. 2019.

Scene memory transformer for embodied agents in long-horizon tasks. In
Proc. of the IEEE Int. Conf. on Computer Vision and Pattern Recognition
. IEEE, 538–547.

Fassold and Takacs (2019)

Hannes Fassold and Barnabas Takacs. 2019.

Towards Automatic Cinematography and Annotation for 360° Video. In
Proc. of the ACM Int. Conf. on Interactive Experiences for TV and Online Video
. ACM, 157–166.

Feichtenhofer et al
.
(2019)

Christoph Feichtenhofer et al
.
2019.

Slowfast networks for video recognition. In
Proc. of the IEEE Int. Conf. on Computer Vision
. IEEE, 6202–6211.

Feichtenhofer et al
.
(2016)

Christoph Feichtenhofer, Axel Pinz, and Andrew Zisserman. 2016.

Convolutional Two-Stream Network Fusion for Video Action Recognition. In
Proc. of the IEEE Int. Conf. on Computer Vision and Pattern Recognition
. IEEE, 1933–1941.

Feng et al
.
(2017)

Yachuang Feng, Yuan Yuan, and Xiaoqiang Lu. 2017.

Learning deep event models for crowd anomaly detection.

Neurocomputing
219 (2017), 548–556.

Fernando et al
.
(2017)

Basura Fernando, Hakan Bilen, Efstratios Gavves, and Stephen Gould. 2017.

Self-supervised video representation learning with odd-one-out networks. In
Proc. of the IEEE Int. Conf. on Computer Vision and Pattern Recognition
. IEEE, 3636–3645.

Gabeur et al
.
(2020)

Valentin Gabeur, Chen Sun, Karteek Alahari, and Cordelia Schmid. 2020.

Multi-modal transformer for video retrieval. In
Proc. of the European Conf. on Computer Vision
. Springer, 214–229.

Gavrilyuk et al
.
(2020)

Kirill Gavrilyuk, Ryan Sanford, Mehrsan Javan, and Cees GM Snoek. 2020.

Actor-transformers for group activity recognition. In
Proc. of the IEEE Int. Conf. on Computer Vision and Pattern Recognition
. IEEE, 839–848.

Girdhar et al
.
(2019)

Rohit Girdhar, Joao Carreira, Carl Doersch, and Andrew Zisserman. 2019.

Video action transformer network. In
Proc. of the IEEE Int. Conf. on Computer Vision and Pattern Recognition
. IEEE, 244–253.

Guo and Lai (2014)

Guodong Guo and Alice Lai. 2014.

A Survey on still-image-based human action recognition.

Pattern Recognition
47, 10 (2014), 3343–3361.

Han et al
.
(2017)

Fei Han, Brian Reily, William Hoff, and Hao Zhang. 2017.

Space-time representation of people based on 3D skeletal data: a review.

Journal of Vision Communication Image Representation
158 (2017), 85–105.

Han et al
.
(2019)

Tengda Han, Weidi Xie, and Andrew Zisserman. 2019.

Video representation learning by dense predictive coding. In
Proc. of the Int. Conf. on Computer Vision Workshops
. IEEE, 1483–1492.

He et al
.
(2016)

Kaiming He et al
.
2016.

Deep residual learning for image recognition. In
Proc. of the IEEE Int. Conf. on Computer Vision and Pattern Recognition
. 770–778.

Hendrycks et al
.
(2020)

Dan Hendrycks, Xiaoyuan Liu, Eric Wallace, Adam Dziedzic, Rishabh Krishnan, and Dawn Song. 2020.

Pretrained transformers improve out-of-distribution robustness.

arXiv:2004.06100
(2020).

Herath et al
.
(2017)

Samitha Herath, Mehrtash Harandi, and Fatih Porikli. 2017.

Going deeper into action recognition: a survey.

Image Vision Computing
60 (2017), 4–21.

Hodgins (2021)

Jessica Hodgins. 2021.

Carnegie Mellon University - CMU Graphics Lab - motion capture library.

http://mocap.cs.cmu.edu/
.

(Accessed on 01/28/2021).

Huang et al
.
(2020)

Yi Huang, Xiaoshan Yang, Junyu Gao, Jitao Sang, and Changsheng Xu. 2020.

Knowledge-Driven Egocentric Multimodal Activity Recognition.

ACM Transactions on Multimedia Computing Communication Applications
16, 4, Article 133 (2020), 133 pages.

Iashin and Rahtu (2020a)

Vladimir Iashin and Esa Rahtu. 2020a.

A better use of audio-visual cues: Dense video captioning with bi-modal transformer.

arXiv:2005.08271
(2020).

Iashin and Rahtu (2020b)

Vladimir Iashin and Esa Rahtu. 2020b.

Multi-modal dense video captioning. In
Proc. of the IEEE Int. Conf. on Computer Vision and Pattern Recognition Workshops
. IEEE, 958–959.

Inc (2020)

ASUSTeK Computer Inc. 2020.

Xtion PRO LIVE| 3D Sensor | ASUS USA.

https://www.asus.com/us/3D-Sensor/Xtion_PRO_LIVE/
.

Accessed: 07-01-2022.

Islam and Iqbal (2020)

Md Mofijul Islam and Tariq Iqbal. 2020.

Hamlet: A hierarchical multimodal attention-based human activity recognition algorithm. In
2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
. IEEE, 10285–10292.

Jaegle et al
.
(2021)

Andrew Jaegle, Felix Gimeno, Andy Brock, Oriol Vinyals, Andrew Zisserman, and Joao Carreira. 2021.

Perceiver: General perception with iterative attention. In
Proc. of the Int. Conf. on Machine Learning
. 4651–4664.

Kalfaoglu et al
.
(2020)

M Kalfaoglu, Sinan Kalkan, and A Aydin Alatan. 2020.

Late temporal modeling in 3D CNN architectures with BERT for action recognition. In
European Conf. on Computer Vision
. Springer, 731–747.

Kay et al
.
(2017)

Will Kay, Joao Carreira, Karen Simonyan, et al
.
2017.

The kinetics human action video dataset.

arXiv:1705.06950
(2017).

Khaire et al
.
(2018)

Pushpajit Khaire, Praveen Kumar, and Javed Imran. 2018.

Combining CNN streams of RGB-D and skeletal data for human activity recognition.

Pattern Recognition Letters
115 (2018), 107 – 116.

Khan et al
.
(2015)

Aftab Khan, Sebastian Mellor, Eugen Berlin, et al
.
2015.

Beyond activity recognition: skill assessment from accelerometer data. In
Proc. of the Int. Joint Conf. on Pervasive and Ubiquitous Computing
. 1155–1166.

Khan et al
.
(2020)

Muhammad Attique Khan, Kashif Javed, Sajid Ali Khan, Tanzila Saba, Usman Habib, Junaid Ali Khan, and Aaqif Afzaal Abbasi. 2020.

Human action recognition using fusion of multiview and deep features: an application to video surveillance.

Multimedia Tools and Applications
(2020), 1–27.

Kim et al
.
(2019)

Dahun Kim, Donghyeon Cho, and In So Kweon. 2019.

Self-supervised video representation learning with space-time cubic puzzles. In
Proc. of the AAAI Conf. on Artificial Intelligence
, Vol. 33. 8545–8552.

Kim et al
.
(2018)

Kyung-Min Kim, Seong-Ho Choi, Jin-Hwa Kim, and Byoung-Tak Zhang. 2018.

Multimodal dual attention memory for video story question answering. In
Proc. of the European Conf. on Computer Vision
. Springer, 673–688.

Kondo (2021)

Satoshi Kondo. 2021.

Lapformer: surgical tool detection in laparoscopic surgical video using transformer architecture.

Computer Methods in Biomechanics and Biomedical Engineering: Imaging & Visualization
9, 3 (2021), 302–307.

Kong et al
.
(2019)

Quan Kong, Ziming Wu, Ziwei Deng, Martin Klinkigt, Bin Tong, and Tomokazu Murakami. 2019.

MMAct: A large-scale dataset for cross modal human action understanding. In
Proc. of the Int. Conf. on Computer Vision
. IEEE, 8658–8667.

Korbar et al
.
(2019)

Bruno Korbar, Du Tran, and Lorenzo Torresani. 2019.

SCSampler: Sampling Salient Clips from Video for Efficient Action Recognition. In
Proc. of the IEEE Int. Conf. on Computer Vision
. IEEE, 6231–6241.

Kuehne et al
.
(2011)

Hilde Kuehne, Hueihan Jhuang, Estibaliz Garrote, Tomaso Poggio, and Thomas Serre. 2011.

HMDB: A Large Video Database for Human Motion Recognition. In
Proc. of the IEEE Int. Conf. on Computer Vision
. IEEE, 2556–2563.

Lane et al
.
(2015)

Nicholas D Lane, Petko Georgiev, and Lorena Qendro. 2015.

Deepear: robust smartphone audio sensing in unconstrained acoustic environments using deep learning. In
Proc. of the Int. Joint Conf. on Pervasive and Ubiquitous Computing
. 283–294.

Lee et al
.
(2020)

Sangho Lee, Youngjae Yu, Gunhee Kim, Thomas Breuel, Jan Kautz, and Yale Song. 2020.

Parameter efficient multimodal transformers for video representation learning.

arXiv:2012.04124
(2020).

Li et al
.
(2020)

Linjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan, Licheng Yu, and Jingjing Liu. 2020.

Hero: Hierarchical encoder for video+ language omni-representation pre-training.

arXiv:2005.00200
(2020).

Li et al
.
(2021)

Shuaicheng Li, Qianggang Cao, Lingbo Liu, Kunlin Yang, Shinan Liu, Jun Hou, and Shuai Yi. 2021.

Groupformer: Group activity recognition with clustered spatial-temporal transformer. In
Proc. of the Int. Conf. on Computer Vision
. IEEE, 13668–13677.

Liciotti et al
.
(2020)

Daniele Liciotti, Michele Bernardini, Luca Romeo, and Emanuele Frontoni. 2020.

A sequential deep learning application for recognising human activities in smart homes.

Neurocomputing
396 (2020), 501–513.

Liu et al
.
(2017)

Chunhui Liu, Yueyu Hu, Yanghao Li, Sijie Song, and Jiaying Liu. 2017.

PKU-MMD: A Large Scale Benchmark for Continuous Multi-Modal Human Action Understanding.

arXiv 1903.11314
(2017).

Liu et al
.
(2019)

Jun Liu, Amir Shahroudy, Mauricio Lisboa Perez, Gang Wang, Ling-Yu Duan, and Alex Kot Chichung. 2019.

NTU RGB+D 120: A Large-Scale Benchmark for 3D Human Activity Understanding.

IEEE Transactions on Pattern Analysis Machine Intelligence
(2019), 2684 – 2701.

Liu and Yuan (2018)

Mengyuan Liu and Junsong Yuan. 2018.

Recognizing Human Actions as the Evolution of Pose Estimation Maps. In
Proc. of the IEEE Int. Conf. on Computer Vision and Pattern Recognition
. IEEE, 1159–1168.

Liu et al
.
(2021)

Song Liu, Haoqi Fan, Shengsheng Qian, Yiru Chen, Wenkui Ding, and Zhongyuan Wang. 2021.

Hit: Hierarchical transformer with momentum contrast for video-text retrieval. In
Proc. of the Int. Conf. on Computer Vision
. IEEE, 11915–11925.

Liu et al
.
(2022)

Ze Liu, Jia Ning, Yue Cao, Yixuan Wei, Zheng Zhang, Stephen Lin, and Han Hu. 2022.

Video swin transformer. In
Proc. of the IEEE Int. Conf. on Computer Vision and Pattern Recognition
. IEEE, 3202–3211.

Long et al
.
(2018)

Xiang Long, Chuang Gan, Gerard de Melo, Xiao Liu, Yandong Li, Fu Li, and Shilei Wen. 2018.

Multimodal Keyless Attention Fusion for Video Classification. In
Proc. of the AAAI Conf. on Artificial Intelligence
. AAAI Press, 7202–7209.

Lu et al
.
(2019)

Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019.

VilBERT: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks.

Advances in neural information processing systems
32 (2019).

Lu et al
.
(2021)

Kevin Lu, Aditya Grover, Pieter Abbeel, and Igor Mordatch. 2021.

Pretrained transformers as universal computation engines.

arXiv:2103.05247
(2021).

Luvizon et al
.
(2018)

D. C. Luvizon, D. Picard, and H. Tabia. 2018.

2D/3D Pose Estimation and Action Recognition Using Multitask Deep Learning. In
Proc. of the IEEE Int. Conf. on Computer Vision and Pattern Recognition
. IEEE, 5137–5146.

Memmesheimer et al
.
(2020)

Raphael Memmesheimer, Nick Theisen, and Dietrich Paulus. 2020.

Gimme Signals: Discriminative signal encoding for multimodal activity recognition.

arXiv:2003.06156
(2020).

Min et al
.
(2016)

Xiongkuo Min, Guangtao Zhai, Ke Gu, and Xiaokang Yang. 2016.

Fixation Prediction through Multimodal Analysis.

ACM Transactions on Multimedia Computing Communication Applications
13, 1, Article 6 (2016), 23 pages.

Mohsen et al
.
(2022)

Farida Mohsen, Hazrat Ali, Nady El Hajj, and Zubair Shah. 2022.

Artificial intelligence-based methods for fusion of electronic health records and imaging data.

Scientific Reports
12, 1 (2022), 1–16.

Nan et al
.
(2019)

Mihai Nan, Alexandra Stefania Ghi

t

,

ă, Alexandru-Florin Gavril, Mihai Trascau, Alexandru Sorici, Bogdan Cramariuc, and Adina Magda Florea. 2019.

Human Action Recognition for Social Robots. In
Proc. of the Int. Conf. on Control Systems and Computer Science
. IEEE, 675–681.

Ni et al
.
(2011)

Bingbing Ni, Gang Wang, and Pierre Moulin. 2011.

RGBD-HuDaAct: A color-depth video database for human daily activity recognition. In
Proc. of the IEEE Int. Conf. on Computer Vision Workshops
. IEEE, 1147–1153.

Nie et al
.
(2020)

Weizhi Nie, Qi Liang, Yixin Wang, Xing Wei, and Yuting Su. 2020.

MMFN: Multimodal Information Fusion Networks for 3D Model Classification and Retrieval.

ACM Transactions on Multimedia Computing Communication Applications
16, 4, Article 131 (2020), 22 pages.

Ofli et al
.
(2013)

F. Ofli, R. Chaudhry, G. Kurillo, R. Vidal, and R. Bajcsy. 2013.

Berkeley MHAD: A comprehensive Multimodal Human Action Database. In
Proc. of the Workshop on Applications of Comp. Vision
. IEEE, Tampa, USA, 53–60.

OpenAI (2023)

OpenAI. 2023.

GPT-4 Technical Report.

arXiv:2303.08774 [cs.CL]

Oreifej and Liu (2013)

Omar Oreifej and Zicheng Liu. 2013.

HON4D: Histogram of Oriented 4D Normals for Activity Recognition from Depth Sequences. In
Proc. of the IEEE Int. Conf. on Computer Vision and Pattern Recognition
. IEEE, 716–723.

Pashevich et al
.
(2021)

Alexander Pashevich, Cordelia Schmid, and Chen Sun. 2021.

Episodic transformer for vision-and-language navigation. In
Proc. of the IEEE Int. Conf. on Computer Vision
. IEEE, 15942–15952.

Patel et al
.
(2018)

Chirag I Patel, Sanjay Garg, Tanish Zaveri, Asim Banerjee, and Ripal Patel. 2018.

Human action recognition using fusion of features for unconstrained video sequences.

Computers & Electrical Engineering
70 (2018), 284–301.

Patrick et al
.
(2021)

Mandela Patrick, Po-Yao Huang, Ishan Misra, et al
.
2021.

Space-time crop & attend: Improving cross-modal video representation learning. In
Proc. of the Int. Conf. on Computer Vision
. IEEE, 10560–10572.

Perez-Rua et al
.
(2019)

Juan-Manuel Perez-Rua, Valentin Vielzeuf, Stephane Pateux, et al
.
2019.

MFAS: Multimodal Fusion Architecture Search. In
Proc. of the IEEE Int. Conf. on Computer Vision and Pattern Recognition
. IEEE, 6966–6975.

Perrett et al
.
(2021)

Toby Perrett, Alessandro Masullo, Tilo Burghardt, Majid Mirmehdi, and Dima Damen. 2021.

Temporal-relational crosstransformers for few-shot action recognition. In
Proc. of the IEEE Int. Conf. on Computer Vision and Pattern Recognition
. IEEE, 475–484.

Prelec et al
.
(2017)

Dražen Prelec, H Sebastian Seung, and John McCoy. 2017.

A solution to the single-question crowd wisdom problem.

Nature
541, 7638 (2017), 532–535.

Purwanto et al
.
(2019)

Didik Purwanto, Rizard Renanda Adhi Pramono, Chen, et al
.
2019.

Extreme low resolution action recognition with spatial-temporal multi-head self-attention and knowledge distillation. In
Proc. of the Int. Conf. on Computer Vision Workshops
. IEEE, 961–969.

Qi et al
.
(2019)

Mengshi Qi, Yunhong Wang, Jie Qin, et al
.
2019.

StagNet: An attentive semantic RNN for group activity and individual action recognition.

IEEE Transactions on Circuits and Systems for Video Technology
30, 2 (2019), 549–565.

Rahman et al
.
(2021)

MD Abdur Rahman, M. Shamim Hossain, Nabil A. Alrajeh, and B. B. Gupta. 2021.

A Multimodal, Multimedia Point-of-Care Deep Learning Framework for COVID-19 Diagnosis.

ACM Transactions on Multimedia Computing Communication Applications
17, 1s, Article 18 (2021), 24 pages.

Ramesh et al
.
(2021)

Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. 2021.

Zero-shot text-to-image generation. In
Int. Conf. on Machine Learning
. PMLR, 8821–8831.

Recasens et al
.
(2021)

Adria Recasens, Pauline Luc, Jean-Baptiste Alayrac, et al
.
2021.

Broaden your views for self-supervised video learning. In
Proc. of the Int. Conf. on Computer Vision
. IEEE, 1255–1265.

Ren et al
.
(2015)

Shaoqing Ren et al
.
2015.

Faster R-CNN: Towards real-time object detection with region proposal networks. In
Proc. of the NeurIPS
, Vol. 28. 91–99.

Roitberg et al
.
(2019)

Alina Roitberg, Tim Pollert, Monica Haurilet, et al
.
2019.

Analysis of Deep Fusion Strategies for Multi-Modal Gesture Recognition. In
Proc. of the IEEE Int. Conf. on Computer Vision and Pattern Recognition Workshops
. IEEE, 198–206.

Rombach et al
.
(2022)

Robin Rombach, Andreas Blattmann, Dominik Lorenz, et al
.
2022.

High-resolution image synthesis with latent diffusion models. In
Proc. of the IEEE Int. Conf. on Computer Vision and Pattern Recognition
. IEEE, 10684–10695.

Selva et al
.
(2022)

Javier Selva, Anders S Johansen, Sergio Escalera, Kamal Nasrollahi, Thomas B Moeslund, and Albert Clapés. 2022.

Video transformers: a survey.

arXiv:2201.05991
(2022).

Shahroudy et al
.
(2016a)

Amir Shahroudy, Jun Liu, Tian-Tsong Ng, and Gang Wang. 2016a.

NTU RGB+D: a large scale dataset for 3D human activity analysis. In
Proc. of the IEEE Int. Conf. on Computer Vision and Pattern Recognition
.

Shahroudy et al
.
(2016b)

Amir Shahroudy, Tian-Tsong Ng, Yihong Gong, and Gang Wang. 2016b.

Deep multimodal feature analysis for action recognition in RGB-D Videos.

arXiv 1603.07120
(2016).

Shaikh and Chai (2021)

Muhammad Bilal Shaikh and Douglas Chai. 2021.

RGB-D data-based action recognition: a review.

Sensors
21, 12 (2021), 4246.

Shaikh et al
.
(2022)

Muhammad Bilal Shaikh, Douglas Chai, Syed Mohammed Shamsul Islam, and Naveed Akhtar. 2022.

MAiVAR: Multimodal Audio-Image and Video Action Recognizer. In
International Conference on Visual Communications and Image Processing (VCIP)
. IEEE, Suzhou, China, 1–5.

Shaikh et al
.
(2024)

Muhammad Bilal Shaikh, Douglas Chai, Syed Mohammed Shamsul Islam, and Naveed Akhtar. 2024.

Multimodal Fusion for Audio-Image and Video Action Recognition.

Neural Computing and Applications
(2024), 1–14.

Shaikh et al
.
(2023)

Muhammad Bilal Shaikh, Douglas Chai, Syed Mohammed Shamsul Islam, and Naveed Akhtar. 2023.

MAiVAR-T: Multimodal Audio-image and Video Action Recognizer using Transformers. In
11th European Workshop on Visual Information Processing (EUVIP)
. 1–6.

Sharif et al
.
(2020)

Muhammad Sharif, Muhammad Attique Khan, Farooq Zahid, et al
.
2020.

Human action recognition: a framework of statistical weighted segmentation and rank correlation-based selection.

Pattern Analysis and Applications
23 (2020), 281–294.

Simonyan and Zisserman (2014)

Karen Simonyan and Andrew Zisserman. 2014.

Two-stream Convolutional Networks for Action Recognition in Videos. In
Proc. of the Int. Conf. on Neural Information Process. Systems (NIPS)
, Vol. 1. MIT Press, 568–576.

Singh et al
.
(2021)

Ankit Singh, Omprakash Chakraborty, Ashutosh Varshney, et al
.
2021.

Semi-supervised action recognition with temporal contrastive learning. In
Proc. of the IEEE Int. Conf. on Computer Vision and Pattern Recognition
. IEEE, 10389–10399.

Song et al
.
(2021)

Xiaolin Song, Sicheng Zhao, Jingyu Yang, et al
.
2021.

Spatio-temporal contrastive domain adaptation for action recognition. In
Proc. of the IEEE Int. Conf. on Computer Vision and Pattern Recognition
. IEEE, 9787–9795.

Soomro et al
.
(2012)

Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah. 2012.

UCF101: A dataset of 101 human actions classes from videos in the wild.

arXiv:1212.0402
(2012).

Su et al
.
(2021)

Rui Su, Qian Yu, and Dong Xu. 2021.

STVGbert: A visual-linguistic transformer based framework for spatio-temporal video grounding. In
Proc. of the Int. Conf. on Computer Vision
. IEEE, 1533–1542.

Sun et al
.
(2019a)

Chen Sun, Fabien Baradel, Kevin Murphy, and Cordelia Schmid. 2019a.

Learning Video Representations using Contrastive Bidirectional Transformer.

arXiv 1906.05743
(2019).

Sun et al
.
(2019b)

Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid. 2019b.

VideoBERT: A joint model for video and language representation learning. In
Proc. of the Int. Conf. on Computer Vision
. IEEE, 7464–7473.

Sun et al
.
(2021)

Chen Sun, Arsha Nagrani, Yonglong Tian, and Cordelia Schmid. 2021.

Composable augmentation encoding for video representation learning. In
Proc. of the Int. Conf. on Computer Vision
. IEEE, 8834–8844.

Sun and Chen (2022)

Han Sun and Yu Chen. 2022.

Real-time Elderly Monitoring for Senior Safety by Lightweight Human Action Recognition. In
Proc. of the IEEE Int. Symposium on Medical Information and Communication Technology
. IEEE, 1–6.

Sun et al
.
(2022)

Zehua Sun, Qiuhong Ke, Hossein Rahmani, et al
.
2022.

Human Action Recognition From Various Data Modalities: A Review.

IEEE Transactions on Pattern Analysis and Machine Intelligence
(2022), 1–20.

Susan et al
.
(2019)

S. Susan, P. Agrawal, M. Mittal, and S. Bansal. 2019.

New shape descriptor in the context of edge continuity.

CAAI Transactions on Intelligence Technology
4 (2019), 101–109.

Tian et al
.
(2018)

Yapeng Tian, Jing Shi, Bochen Li, Zhiyao Duan, and Chenliang Xu. 2018.

Audio-visual event localization in unconstrained videos. In
Proc. of the European Conf. on Computer Vision
. Springer, 247–263.

Tingting et al
.
(2019)

Y. Tingting, W. Junqian, W. Lintai, and X. Yong. 2019.

Three-stage network for age estimation.

CAAI Transactions on Intelligence Technology
4 (2019), 122–126.

Tran et al
.
(2018)

Du Tran, Heng Wang, Lorenzo Torresani, et al
.
2018.

A closer look at spatiotemporal convolutions for action recognition. In
Proc. of the IEEE Int. Conf. on Computer Vision and Pattern Recognition
. IEEE, 6450–6459.

Ullah et al
.
(2019)

Amin Ullah, Khan Muhammad, Ijaz Ul Haq, and Sung Wook Baik. 2019.

Action recognition using optimized deep autoencoder and CNN for surveillance data streams of non-stationary environments.

Future Generation Computer Systems
96 (2019), 386–397.

Vaezi Joze et al
.
(2020)

Hamid Reza Vaezi Joze, Amirreza Shaban, Michael L. Iuzzolino, and Kazuhito Koishida. 2020.

MMTM: Multimodal Transfer Module for CNN Fusion. In
Proc. of the IEEE Int. Conf. on Computer Vision and Pattern Recognition
. IEEE, Virtual, 13289–13299.

Vaswani et al
.
(2017)

Ashish Vaswani, Noam Shazeer, Niki Parmar, et al
.
2017.

Attention is all you need. In
Advances in Neural Information Processing Systems
, Vol. 30.

Vepakomma et al
.
(2015)

Praneeth Vepakomma, Debraj De, Sajal K Das, and Shekhar Bhansali. 2015.

A-Wristocracy: Deep learning on wrist-worn sensing for recognition of user complex activities. In
Proc. of the IEEE Int. Conf. on Wearable and Implantable Body Sensor Networks (BSN)
. IEEE, 1–6.

Wang et al
.
(2017)

Jindong Wang, Yiqiang Chen, Shuji Hao, et al
.
2017.

Balanced distribution adaptation for transfer learning. In
Proceddings of the IEEE Int. Conf. on Data Mining
. IEEE, 1129–1134.

Wang et al
.
(2019a)

Jiangliu Wang, Jianbo Jiao, Linchao Bao, et al
.
2019a.

Self-supervised spatio-temporal representation learning for videos by predicting motion and appearance statistics. In
Proc. of the IEEE Int. Conf. on Computer Vision and Pattern Recognition
. IEEE, 4006–4015.

Wang et al
.
(2012)

Jiang Wang, Zicheng Liu, Ying Wu, and Junsong Yuan. 2012.

Mining actionlet ensemble for action recognition with depth cameras. In
Proc. of the IEEE Int. Conf. on Computer Vision and Pattern Recognition
. IEEE, 1290–1297.

Wang et al
.
(2014a)

Jiang Wang, Xiaohan Nie, Yin Xia, Ying Wu, and Song-Chun Zhu. 2014a.

Cross-view action modeling, learning and recognition. In
Proc. of the IEEE Int. Conf. on Computer Vision and Pattern Recognition
. IEEE, 2649–2656.

Wang et al
.
(2014b)

Jiang Wang, Xiaohan Nie, Yin Xia, Ying Wu, and Song-Chun Zhu. 2014b.

Cross-view action modelling, learning and recognition. In
Proc. of the IEEE Int. Conf. on Computer Vision and Pattern Recognition
. Columbus, USA, 2649–2656.

Wang et al
.
(2020a)

Jingdong Wang, Ke Sun, Tianheng Cheng, et al
.
2020a.

Deep high-resolution representation learning for visual recognition.

IEEE Transactions on Pattern Analysis and Machine Intelligence
43, 10 (2020), 3349–3364.

Wang et al
.
(2016)

Limin Wang et al
.
2016.

Temporal segment networks: Towards good practices for deep action recognition. In
Proc. of the European Conf. on Computer Vision
. Springer, 20–36.

Wang et al
.
(2018)

Pichao Wang, Wanqing Li, Philip Ogunbona, Jun Wan, and Sergio Escalera. 2018.

RGB-D-based human motion recognition with deep learning: a survey.

Computing Vision Image Understanding
171 (2018), 118–139.

Wang et al
.
(2020b)

Weiyao Wang, Du Tran, and Matt Feiszli. 2020b.

What makes training multi-modal classification networks hard?. In
Proc. of the IEEE Int. Conf. on Computer Vision and Pattern Recognition
. IEEE, 12695–12705.

Wang (2021)

Yang Wang. 2021.

Survey on Deep Multi-Modal Data Analytics: Collaboration, Rivalry, and Fusion.

ACM Transactions on Multimedia Computing Communication Applications
17, 1s, Article 10 (2021), 25 pages.

Wang et al
.
(2019b)

Zhen Wang, Shixian Luo, He Sun, et al
.
2019b.

An efficient non-local attention network for video-based person re-identification. In
Proc. of the Int. Conf. on Information Technology: IoT and Smart City
. ACM, 212–217.

Wang et al
.
(2019c)

Zihao W. Wang, Vibhav Vineet, Francesco Pittaluga, et al
.
2019c.

Privacy-Preserving Action Recognition Using Coded Aperture Videos. In
Proc. of the IEEE Int. Conf. on Computer Vision and Pattern Recognition Workshops
. 1–10.

Wu et al
.
(2021)

Kan Wu, Houwen Peng, Minghao Chen, Jianlong Fu, and Hongyang Chao. 2021.

Rethinking and improving relative position encoding for vision transformer. In
Proc. of the Int. Conf. on Computer Vision
. IEEE, 10033–10041.

Xia et al
.
(2012)

L. Xia, C.C. Chen, and JK Aggarwal. 2012.

View Invariant Human Action Recognition using Histograms of 3D Joints. In
Proc. of the IEEE Int. Conf. on Computer Vision and Pattern Recognition
. IEEE, 20–27.

Xie et al
.
(2017)

Saining Xie, Chen Sun, Jonathan Huang, Zhuowen Tu, and Kevin Murphy. 2017.

Rethinking spatiotemporal feature learning for video understanding.

arXiv:1712.04851
1, 2 (2017), 5.

Xu et al
.
(2021)

Tong Xu, Peilun Zhou, Linkang Hu, Xiangnan He, Yao Hu, and Enhong Chen. 2021.

Socializing the Videos: A Multimodal Approach for Social Relation Recognition.

ACM Transactions on Multimedia Computing Communication Applications
17, 1, Article 23 (2021), 23 pages.

Yang et al
.
(2015)

Lin Yang, Longyu Zhang, Haiwei Dong, Abdulhameed Alelaiwi, and Abdulmotaleb El Saddik. 2015.

Evaluating and improving the depth accuracy of Kinect for Windows v2.

IEEE Sensors
15 (2015), 4275–4285.

Yin et al
.
(2022)

Guanghao Yin, Shouqian Sun, Dian Yu, Dejian Li, and Kejun Zhang. 2022.

A Multimodal Framework for Large-Scale Emotion Recognition by Fusing Music and Electrodermal Activity Signals.

ACM Transactions on Multimedia Computing Communication Applications
18, 3, Article 78 (2022), 23 pages.

Yu et al
.
(2021)

Bingyao Yu, Wanhua Li, Xiu Li, Jiwen Lu, and Jie Zhou. 2021.

Frequency-aware spatiotemporal transformers for video inpainting detection. In
Proc. of the Int. Conf. on Computer Vision
. IEEE, 8188–8197.

Yun et al
.
(2012)

Kiwon Yun, Jean Honorio, Debaleena Chattopadhyay, et al
.
2012.

Two-person Interaction Detection Using Body-Pose Features and Multiple Instance Learning. In
Proc. of the IEEE Int. Conf. on Computer Vision and Pattern Recognition Workshops
. IEEE, 28–35.

Zahin et al
.
(2019)

Abrar Zahin, Rose Qingyang Hu, et al
.
2019.

Sensor-based human activity recognition for smart healthcare: a semi-supervised machine learning. In
Proceedins of the Int. Conf. on Artificial Intelligence for Communications and Networks
. Springer, 450–472.

Zbontar et al
.
(2021)

Jure Zbontar, Li Jing, Ishan Misra, Yann LeCun, and Stéphane Deny. 2021.

Barlow twins: self-supervised learning via redundancy reduction. In
Int. Conf. on Machine Learning
. PMLR, 12310–12320.

Zha et al
.
(2021)

Xuefan Zha, Wentao Zhu, Lv Xun, Sen Yang, and Ji Liu. 2021.

Shifted chunk transformer for spatio-temporal representational learning.

Advances in Neural Information Processing Systems
34 (2021), 11384–11396.

Zhang et al
.
(2022)

Chongzhi Zhang, Mingyuan Zhang, Zhang, et al
.
2022.

Delving deep into the generalization of vision transformers under distribution shifts. In
Proc. of the IEEE Int. Conf. on Computer Vision and Pattern Recognition
. IEEE, 7277–7286.

Zhang et al
.
(2019b)

Hong Bo Zhang, Yi Xiang Zhang, Bineng Zhong, et al
.
2019b.

A Comprehensive Survey of Vision-Based Human Action Recognition Methods.

Sensors
19, 5 (2019), 1005.

Zhang et al
.
(2016)

Jing Zhang, Wanqing Li, Philip O. Ogunbona, et al
.
2016.

RGB-D-based action recognition datasets: a survey.

Pattern Recognition
60 (2016), 86–105.

Zhang et al
.
(2021)

Mingxing Zhang, Yang Yang, Xinghan Chen, et al
.
2021.

Multi-stage aggregated transformer network for temporal language localization in videos. In
Proc. of the IEEE Int. Conf. on Computer Vision and Pattern Recognition
. IEEE, 12669–12678.

Zhang et al
.
(2019a)

Wei Zhang, Ting Yao, Shiai Zhu, and Abdulmotaleb El Saddik. 2019a.

Deep learning–based Multimedia Analytics: A Review.

ACM Transactions on Multimedia Computing Communication Applications
15, 1s, Article 2 (2019), 26 pages.

Zhang et al
.
(2017)

Z. Zhang, X. Ma, R. Song, X. Rong, X. Tian, G. Tian, and Y. Li. 2017.

Deep learning-based human action recognition: a survey. In
Chinese Automation Congress (CAC)
. IEEE, 3780–3785.

Zhu et al
.
(2016)

Fan Zhu, Ling Shao, Jin Xie, and Yi Fang. 2016.

From handcrafted to learned representations for human action recognition: a survey.

Image Vision Computing
55 (2016), 42–52.

Zhu et al
.
(2016)

G. Zhu, L. Zhang, L. Mei, Jie Shao, Juan Song, and Peiyi Shen. 2016.

Large-scale Isolated Gesture Recognition using pyramidal 3D convolutional networks. In
Proc. of the Int. Conf. on Pattern Recognition
. IEEE, 19–24.

Zhu et al
.
(2018)

Jiagang Zhu, Wei Zou, Liang Xu, et al
.
2018.

Action Machine: Rethinking Action Recognition in Trimmed Videos.

arXiv 1812.05770
(2018).

Zhu and Yang (2020)

Linchao Zhu and Yi Yang. 2020.

ActBERT: Learning global-local video-text representations. In
Proc. of the IEEE Int. Conf. on Computer Vision and Pattern Recognition
. IEEE, 8746–8755.

Zhuang et al
.
(2017)

Bohan Zhuang, Lingqiao Liu, Yao Li, Chunhua Shen, and Ian Reid. 2017.

Attend in Groups: A Weakly-Supervised Deep Learning Framework for Learning From Web Data. In
Proc. of the IEEE Int. Conf. on Computer Vision and Pattern Recognition
. IEEE, 1878–1887.

Zlatintsi et al
.
(2020)

Athanasia Zlatintsi, AC Dometios, Nikolaos Kardaris, et al
.
2020.

I-Support: a robotic platform of an assistive bathing robot for the elderly population.

Robotics and Autonomous Systems
126 (2020), 103451.

◄

Feeling
lucky?

Conversion
report

Report
an issue

View original
on arXiv
►

Copyright

Privacy Policy

Generated by
L
a
T
e
XML
oxide
</reference>

<statements>
1. A human-action-recognition survey states that multimodal data leads to superior performance compared with a single modality, and that fusion aims to achieve better accuracy than a single modality
2. It lists benefits such as enhanced signal-to-noise ratio, improved confidence, increased robustness, enhanced resolution and better precision
3. It also notes that Transformers are naturally suited to multi-stream fusion because of non-specific embedding and dynamic attention, but that they require substantial computation and memory and are constrained by scarce large-scale multimodal datasets
4. The survey emphasizes that the choice of fusion strategy depends on data characteristics and task requirements and needs empirical evaluation
5. Fusion improves benchmark action recognition and can generate coaching text in research settings
</statements>

Begin the assessment now. Output only the JSON list, without any conversational text or explanations.