Please use this identifier to cite or link to this item:
http://repo.lib.jfn.ac.lk/ujrr/handle/123456789/12903Full metadata record
| DC Field | Value | Language |
|---|---|---|
| dc.contributor.author | Kugarajeevan, J. | - |
| dc.contributor.author | Kokul, T. | - |
| dc.contributor.author | Ramanan, A. | - |
| dc.contributor.author | Fernando, S. | - |
| dc.date.accessioned | 2026-08-24T05:26:44Z | - |
| dc.date.available | 2026-08-24T05:26:44Z | - |
| dc.date.issued | 2024 | - |
| dc.identifier.uri | http://repo.lib.jfn.ac.lk/ujrr/handle/123456789/12903 | - |
| dc.description.abstract | Recent approaches in single-object tracking leverage Transformer-based architectures to achieve state-of-the-art performance by utilizing the attention mechanism between target templates and search region patches. However, most existing trackers fail to capture the temporal information of the target object present in the sequences of frames, making them unsuitable for capturing changes in the target’s appearance over time. This leads to suboptimal performance in real-world scenarios where the target’s appearance can change significantly due to variations such as occlusions, deformations, and abrupt movements. To overcome this limitation, a Transformer-based tracking framework is introduced that simultaneously captures the spatial and temporal features of the target object and utilizes that knowledge to accurately locate the target. This is achieved by adapting the standard Transformer encoder architecture, enabling the direct extraction of spatiotemporal features of the target object from a sequence of frames. As the first step, the sequences of template frames and search region are partitioned into non-overlapping spatiotemporal patches, which are then tokenized and processed by the Transformer encoder to capture spatiotemporal features. Additionally, the Transformer encoder is initialized with spatiotemporal masked autoencoders, which are trained to capture the relationship between patches through self-supervised learning. Extensive experiments on the GOT-10k benchmark dataset demonstrate that our tracker surpasses current state-of-the-art methods, achieving an average overlap of 70.9%, a success rate of 82.4% at a 0.5 threshold, and a success rate of 66.3% at a 0.75 threshold on test data, while maintaining real-time processing speed of 47 frames per second on an NVIDIA P100 GPU. | en_US |
| dc.language.iso | en | en_US |
| dc.publisher | IEEE | en_US |
| dc.subject | Transformer tracking | en_US |
| dc.subject | Spatiotemporal tracking | en_US |
| dc.subject | Visual object tracking | en_US |
| dc.title | Transformer Tracking Using Spatiotemporal Features | en_US |
| dc.type | Conference paper | en_US |
| Appears in Collections: | Computer Science | |
Files in This Item:
| File | Description | Size | Format | |
|---|---|---|---|---|
| Transformer Tracking Using Spatiotemporal Features (Abstract).pdf | 206.72 kB | Adobe PDF | View/Open |
Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.