Ship wake segmentation network based on attention mechanism
-
摘要: 针对传统深度学习语义分割网络难以高精度分割舰船尾迹的问题,提出一种基于通道先验卷积注意力(channel prior convolutional attention, CPCA)机制的VGG-UNet语义分割网络模型。首先,通过资料调查和目视解译构建光学遥感图像舰船尾迹数据集,通过数据增强技术扩充数据集;其次,改进U-Net网络架构,在编码器部分使用视觉几何组(visual geometry group, VGG)主干特征提取网络,将CPCA机制引入VGG-UNet网络模型中的跳跃连接部分,并采用迁移学习策略提升特征获取能力;最后,对改进的CPCA VGG-UNet网络模型进行训练,并进行对比实验和消融实验。结果表明,改进后网络模型的平均交并比Im、平均召回率Rm和平均像素级准确率Pm评价指标分别为87.68%、92.67%和91.58%,均优于U-Net、VGG-UNet和Res-UNet网络模型。本研究所提网络模型具有更高的分割精确度,为更加完整地分割舰船尾迹提供了新思路。Abstract: Aiming at the problem that traditional deep learning semantic segmentation networks are difficult to achieve high-precision segmentation of ship wakes, a VGG-UNet semantic segmentation network model based on the channel prior convolutional attention (CPCA) was proposed. First, a ship wake dataset of optical remote sensing images was constructed through a literature review and visual interpretation, and the dataset was expanded through data augmentation techniques. Second, the U-Net network architecture was improved by incorporating a visual geometry group (VGG) backbone for feature extraction in the encoder section, integrating the CPCA into the skip connection part of the VGG-UNet model, and adopting transfer learning strategies to enhance feature acquisition capabilities. Finally, the improved VGG-UNet was trained, and comparison experiments and ablation experiments were conducted. The experimental results demonstrate that the improved network model achieves 87.68%, 92.67%, and 91.58% on the three evaluation metrics of mean intersection over union Im, mean recall Rm and mean pixel accuracy Pm, respectively, and all its evaluation metrics are superior to those of the U-Net, VGG-UNet, and Res-UNet network models. The proposed network model exhibits higher segmentation accuracy, providing a new approach for more complete segmentation of ship wakes.
-
Key words:
- optical remote sensing images /
- ship wakes /
- semantic segmentation /
- attention mechanism
-
表 1 数据增强效果示意图
Table 1. Illustration of data augmentation effects
图像 原始图像 旋转90° 旋转180° 旋转270° 垂直翻转 水平翻转 原始图像 





标签图像 





表 2 网络训练参数
Table 2. Network training parameters
参数 数值/类别 训练轮次 120 批次大小 2 初始学习率 0.0001 学习率衰减优化器 ADAM 表 3 混淆矩阵
Table 3. Confusion matrix
预测结果 真实情况 舰船尾迹 背景 舰船尾迹 TP FP 背景 FN TN 表 4 对比实验数据
Table 4. Comparative experimental data
网络模型 Im/% Rm/% Pm/% U-Net 81.32 88.37 88.14 VGG-UNet 84.51 91.18 89.73 Res-UNet 83.90 90.27 88.51 CPCA VGG-UNet 87.68 92.67 91.58 表 5 可视化结果
Table 5. Visualized results
图像序号 原始图像 标签图像 U-Net VGG-UNet Res-UNet CPCA VGG-UNet 1 





2 





3 





表 6 消融实验数据
Table 6. Ablation experiment data
编号 CPCA 预训练权重 Im/% Rm/% Pm/% 1 — — 83.14 89.55 88.27 2 — √ 84.51 91.18 89.73 3 √ — 86.39 91.64 90.04 4 √ √ 87.68 92.67 91.58 -
[1] 李国选, 刘凤霞. 海洋命运共同体的科技规制探赜[J] . 自然辩证法研究, 2024, 40(3): 97 − 103. doi: 10.19484/j.cnki.1000-8934.2024.03.014 [2] 林明森, 何贤强, 贾永君, 等. 中国海洋卫星遥感技术进展[J] . 海洋学报, 2019, 41(10): 99 − 112. doi: 10.3969/j.issn.0253-4193.2019.10.007 [3] 李岩. 基于“高分五号”卫星红外影像的舰船尾迹特征分析[J] . 航天返回与遥感, 2020, 41(5): 102 − 109. doi: 10.3969/j.issn.1009-8518.2020.05.012 [4] Tings B, Pleskachevsky A, Wiehle S. Comparison of detectability of ship wake components between C-Band and X-Band synthetic aperture radar sensors operating under different slant ranges[J] . ISPRS Journal of Photogrammetry and Remote Sensing, 2023, 196: 306 − 324. doi: 10.1016/j.isprsjprs.2022.12.008 [5] Karakuş O, Rizaev I, Achim A. Ship wake detection in SAR images via sparse regularization[J] . IEEE Transactions on Geoscience and Remote Sensing, 2020, 58(3): 1665 − 1677. doi: 10.1109/TGRS.2019.2947360 [6] Liu Y F, Zhao J, Qin Y. A novel technique for ship wake detection from optical images[J] . Remote Sensing of Environment, 2021, 258: 112375. doi: 10.1016/j.rse.2021.112375 [7] 冯长峰, 王春平, 付强, 等. 基于深度学习的光学遥感图像目标检测综述[J] . 激光与红外, 2023, 53(9): 1309 − 1319. doi: 10.3969/j.issn.1001-5078.2023.09.002 [8] 吴荣峰, 唐希源. 一种改进的Mask R-CNN卫星影像船舶尾迹检测方法[J] . 智能计算机与应用, 2022, 12(2): 73 − 78 doi: 10.3969/j.issn.2095-2163.2022.02.015 [9] Ding K Y, Yang J F, Lin H, et al. Towards real-time detection of ships and wakes with lightweight deep learning model in Gaofen-3 SAR images[J] . Remote Sensing of Environment, 2023, 284: 113345. doi: 10.1016/j.rse.2022.113345 [10] Redmon J, Divvala S, Girshick R, et al. You only look once: unified, real-time object detection[C] //2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Las Vegas: IEEE, 2016: 779 − 788. [11] 张鑫, 姚庆安, 赵健, 等. 全卷积神经网络图像语义分割方法综述[J] . 计算机工程与应用, 2022, 58(8): 45 − 57. [12] Shelhamer E, Long J, Darrell T. Fully convolutional networks for semantic segmentation[J] . IEEE Transactions on Pattern Analysis and Machine Intelligence, 2017, 39(4): 640 − 651. doi: 10.1109/TPAMI.2016.2572683 [13] Simonyan K, Zisserman A. Very deep convolutional networks for large-scale image recognition[PP/OL] . V5. (2015-04-10)[2024-01-06] . https://doi.org/10.48550/arXiv.1409.1556. [14] Ronneberger O, Fischer P, Brox T. U-Net: convolutional networks for biomedical image segmentation[C] //The 18th International Conference on Medical Image Computing and Computer-Assisted Intervention – MICCAI 2015. Munich: Springer, 2015: 234 − 241. [15] Badrinarayanan V, Kendall A, Cipolla R. SegNet: a deep convolutional encoder-decoder architecture for image segmentation[J] . IEEE Transactions on Pattern Analysis and Machine Intelligence, 2017, 39(12): 2481 − 2495. doi: 10.1109/TPAMI.2016.2644615 [16] Huang H J, Chen Z G, Zou Y, et al. Channel prior convolutional attention for medical image segmentation[J] . Computers in Biology and Medicine, 2024, 178: 108784. [17] Woo S, Park J, Lee J Y, et al. CBAM: convolutional block attention module[C] //Proceedings of 15th European Conference on Computer Vision – ECCV 2018. Munich: Springer, 2018: 3 − 19. -
下载: