권호기사보기
| 기사명 | 저자명 | 페이지 | 원문 | 기사목차 |
|---|
결과 내 검색
동의어 포함
Title Page 2
Contents 5
Abstract 9
Chapter 1. Introduction 10
1.1. Motivation 10
1.2. Overview of the Proposed Methods 13
Chapter 2. Related Works 15
2.1. CNNs 15
2.2. RNNs 18
2.3. GNNs 19
2.4. Transformers 20
Chapter 3. Proposed Methods 22
3.1. State Space Model (SSM) 22
3.2. Proposed Methods 26
Chapter 4. Experiments 33
4.1. Dataset and Evaluation Metrics 33
4.2. Evaluation on Benchmarks 34
4.3. Ablation studies 35
Chapter 5. Conclusion 39
5.1. Conclusion 39
5.2. Limitations and Future work 40
Bibliography 41
Figure 3.1. The architecture overview of proposed methods 26
Figure 3.2. Diagrams of the Embedding, Stem, and Branch modules 27
Figure 3.3. Process of extracting spatiotemporal features using the pretrained video encoder 29
Figure 3.4. Diagrams of the Embedding, Stem, and Branch modules 30
Temporal Action Localization (TAL) is crucial for understanding actions within videos by classifying them and determining their temporal boundaries. Traditional deep learning approaches using CNNs, RNNs, GCNs, and Transformers have made significant strides yet often struggle with capturing long-range dependencies (LRD) and efficiently processing extensive video sequences. This thesis derives insights from these methods to enhance TAL performance, specifically by leveraging a combination of Feature aggregated Bi- S6 block design, Dual Bi-S6 structure, and Recurrent mechanism.
Our approach involves three key components. First, the Feature aggregated Bi-S6 block uses multiple Conv1D layers with various kernel sizes in parallel, summing their outputs to capture local contexts of different ranges. This aggregated result is then fed into the Bi-directional S6 (Bi-S6), enhancing its capacity to model complex features. Sec-ond, the Dual Bi-S6 structure employs two parallel Feature Aggregated Bi-S6 blocks: one (TFA-Bi-S6) processes the temporal dimension, and the other (C-Bi-S6) processes the channel dimension. The outputs are then combined using point-wise multiplication, effec-tively integrating temporal dependencies of spatiotemporal features and spatiotemporal dependencies of temporal features, thereby making TAL more robust. Third, the Recur-rent mechanism applies the Dual Bi-S6 structure recursively r times in a residual manner, allowing the model to iteratively refine its representation of long-range dependencies.
The proposed architecture, inspired by ActionFormer and ActionMamba, includes a Pretrained video encoder that extracts spatiotemporal features for each clip of a video. It features a Backbone that captures dependencies and extracts features at various temporal resolutions from the sequence data containing spatiotemporal features of each clip. The architecture also includes a simple post-processing Neck for handling multiple temporal resolutions and a Head that classifies actions and performs regression on video segments using the post-processed results.
Extensive evaluation on benchmark datasets such as THUMOS-14, ActivityNet, Fine- Action, and HACS validates our approach, with our models outperforming existing state-of-
the-art solutions. We achieved remarkable mean Average Precision (mAP) scores of 74.2% on THUMOS-14, 42.9% on ActivityNet, 29.6% on FineAction, and 45.8% on HACS.*표시는 필수 입력사항입니다.
| 전화번호 |
|---|
| 기사명 | 저자명 | 페이지 | 원문 | 기사목차 |
|---|
| 번호 | 발행일자 | 권호명 | 제본정보 | 자료실 | 원문 | 신청 페이지 |
|---|
도서위치안내: / 서가번호:
우편복사 목록담기를 완료하였습니다.
*표시는 필수 입력사항입니다.
저장 되었습니다.