Ship type classification from AIS-labelled Sentinel-2 imagery
Ocean Engineering, vol.363, 2026 (SCI-Expanded, Scopus)
- Publication Type: Article / Article
- Volume: 363
- Publication Date: 2026
- Doi Number: 10.1016/j.oceaneng.2026.126704
- Journal Name: Ocean Engineering
- Journal Indexes: Science Citation Index Expanded (SCI-EXPANDED), Scopus, Applied Science & Technology Source, Compendex, Environment Index, Geobase, ICONDA Bibliographic, INSPEC, The International Construction Database (ICONDA), Academic Search Ultimate (EBSCO), Engineering Source (EBSCO)
- Keywords: AIS, Maritime surveillance, Multispectral imagery, Sentinel-2, Ship type classification
- Hacettepe University Affiliated: Yes
Abstract
Ship type classification is important for maritime surveillance and safety. Although the Automatic Identification System (AIS) provides vessel identity and type information, AIS-based monitoring can be affected by coverage gaps, transponder deactivation, spoofing, and limited independent visual verification. This study therefore investigates AIS-labelled optical ship type classification under cooperative-target assumptions using Sentinel-2 imagery. The objectives are to construct an AIS-linked Sentinel-2 ship-patch dataset through a reproducible computational pipeline and to establish a transparent ResNet baseline under different label granularities. Approximately 75,000 ship-centred patches were generated from Sentinel-2 images acquired between 2022 and 2024 and automatically labelled by spatio-temporal association with open Danish AIS records. In stratified 10-fold cross-validation, the original 6-class scheme (Cargo, Tanker, Fishing, Passenger, Sailing, Pleasure) achieved 0.64 macro-averaged F1-score and 0.67 accuracy with ResNet-18. Merging visually ambiguous classes into a 4-class taxonomy increased performance to 0.80 macro-averaged F1-score and 0.86 accuracy, while ResNet-34 produced comparable results, indicating that label design matters more than model depth at Sentinel-2 resolution. To demonstrate the effectiveness of the proposed pipeline and assess the generalisability of the Denmark-trained models across different geographical settings, an external replication study was conducted in South Florida. The Denmark-trained models achieved an accuracy of 0.32 and a macro-averaged F1 score of 0.29 in the 6-class setting, and an accuracy of 0.36 and a macro-averaged F1 score of 0.29 in the 4-class setting. These results suggest that enhancing model generalisability across different regions may require larger models and more geographically diverse training datasets.