R-CNN系列 Region-CNN的缩写,主要用于目标检测。 来自 2014 年 CVPR 论文“Rich feature hierarchies for accurate object detection and semantic segmentation” 在 Pascal VOC 2012 的数据集上,能够将目标检测的验证指标 mAP 提升到 53.3%,这相对于之前最好的结果提升了整整 30% 采用在ImageNet上已经训练好的模型,然后在PASCAL VOC数据集上进行 fine-tune 参考:Ross B. Girshick, Jeff Donahue, Trevor Darrell, Jitendra Malik.
参考:Ross B. Girshick, Jeff Donahue, Trevor Darrell, Jitendra Malik. Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation. CVPR 2014: 580-587
区域划分:给定一张输入图片,从图片中提取2000个类别独立的候选区域,R-CNN 采用的是 Selective Search 算法
特征提取:对于每个区域利用CNN抽取一个固定长度的特征向量, R-CNN 使用的是 Alexnet
目标分类:再对每个区域利用SVM进行目标分类
边框回归:BoundingboxRegression(Bbox回归)进行边框坐标偏移
优化和调整
核心思想:图像中物体可能存在的区域应该有某些相似性或者连续性的,选择搜索基于上面这一想法采用子区域合并的方法提取 bounding boxes候选边界框。
算法步骤:
核心思想:通过平移和缩放方法对物体边框进行调整和修正。
mAP:mean Average Precision,是多标签图像分类任务中的评价指标。AP衡量的是学出来的模型在给定类别上的好坏,而mAP衡量的是学出的模型在所有类别上的好坏。
SPPnet (Spatial Pyramid Pooling):空间金字塔网络,R-CNN主要问题:每个Proposal独立提取CNN features,分步训练。
参考:Kaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun. Spatial Pyramid Pooling in Deep Convolutional Networks for Visual Recognition. IEEE Trans. Pattern Anal. Mach. Intell. 37(9): 1904-1916 (2015)
ROI Pooling层:将每个候选区域均匀分成M×N块,对每块进行max pooling,将特征图上大小不一的候选区域转变为大小统一的数据,送入下一层。
性能对比
效率对比
RPN(Region Proposal Network):使用全卷积神经网络来生成区域建议(Region proposal),替代之前的Selective search。
参考:Shaoqing Ren, Kaiming He, Ross B. Girshick, Jian Sun. Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks. IEEE Trans. Pattern Anal. Mach. Intell. 39(6): 1137-1149 (2017)
Faster R-CNN训练方式
Alternating training
Approximate joint training
参考:Joseph Redmon, Santosh Kumar Divvala, Ross B. Girshick, Ali Farhadi. You Only Look Once: Unified, Real-Time Object Detection. CVPR 2016: 779-788
The YOLO Detection System. Processing images with YOLO is simple and straightforward. Our system (1) resizes the input image to 448 ×448, (2) runs a single convolutional network on the image, and (3) thresholds the resulting detections by the model’s confidence.
The Model. Our system models detection as a regression problem. It divides the image into an S×S grid and for each grid cell predicts B bounding boxes, confidence for those boxes, and C class probabilities. These predictions are encoded as an S×S×(B∗5+ C) tensor.
网络结构:24个卷积层和2个全连接层
The Architecture. Our detection network has 24 convolutional layers followed by 2 fully connected layers. Alternating 1×1 convolutional layers reduce the features space from preceding layers. We pretrain the convolutional layers on the ImageNet classification task at half the resolution (224 ×224 input image) and then double the resolution for detection.
参考:Joseph Redmon, Ali Farhadi. YOLO9000: Better, Faster, Stronger. CVPR 2017: 6517-6525
性能分析
