. 2022 May 31;22(11):4187.

doi: 10.3390/s22114187.

Pear Recognition in an Orchard from 3D Stereo Camera Datasets to Develop a Fruit Picking Mechanism Using Mask R-CNN

Siyu Pan¹, Tofael Ahamed²

Affiliations

¹ Graduate School of Science and Technology, University of Tsukuba, 1-1-1 Tennodai, Tsukuba 305-8577, Japan.
² Faculty of Life and Environmental Sciences, University of Tsukuba, 1-1-1 Tennodai, Tsukuba 305-8577, Japan.

PMID: 35684807
PMCID: PMC9185418
DOI: 10.3390/s22114187

Pear Recognition in an Orchard from 3D Stereo Camera Datasets to Develop a Fruit Picking Mechanism Using Mask R-CNN

Siyu Pan et al. Sensors (Basel). 2022.

. 2022 May 31;22(11):4187.

doi: 10.3390/s22114187.

Authors

Siyu Pan¹, Tofael Ahamed²

Affiliations

¹ Graduate School of Science and Technology, University of Tsukuba, 1-1-1 Tennodai, Tsukuba 305-8577, Japan.
² Faculty of Life and Environmental Sciences, University of Tsukuba, 1-1-1 Tennodai, Tsukuba 305-8577, Japan.

PMID: 35684807
PMCID: PMC9185418
DOI: 10.3390/s22114187

Abstract

In orchard fruit picking systems for pears, the challenge is to identify the full shape of the soft fruit to avoid injuries while using robotic or automatic picking systems. Advancements in computer vision have brought the potential to train for different shapes and sizes of fruit using deep learning algorithms. In this research, a fruit recognition method for robotic systems was developed to identify pears in a complex orchard environment using a 3D stereo camera combined with Mask Region-Convolutional Neural Networks (Mask R-CNN) deep learning technology to obtain targets. This experiment used 9054 RGBA original images (3018 original images and 6036 augmented images) to create a dataset divided into a training, validation, and testing sets. Furthermore, we collected the dataset under different lighting conditions at different times which were high-light (9-10 am) and low-light (6-7 pm) conditions at JST, Tokyo Time, August 2021 (summertime) to prepare training, validation, and test datasets at a ratio of 6:3:1. All the images were taken by a 3D stereo camera which included PERFORMANCE, QUALITY, and ULTRA models. We used the PERFORMANCE model to capture images to make the datasets; the camera on the left generated depth images and the camera on the right generated the original images. In this research, we also compared the performance of different types with the R-CNN model (Mask R-CNN and Faster R-CNN); the mean Average Precisions (mAP) of Mask R-CNN and Faster R-CNN were compared in the same datasets with the same ratio. Each epoch in Mask R-CNN was set at 500 steps with total 80 epochs. And Faster R-CNN was set at 40,000 steps for training. For the recognition of pears, the Mask R-CNN, had the mAPs of 95.22% for validation set and 99.45% was observed for the testing set. On the other hand, mAPs were observed 87.9% in the validation set and 87.52% in the testing set using Faster R-CNN. The different models using the same dataset had differences in performance in gathering clustered pears and individual pear situations. Mask R-CNN outperformed Faster R-CNN when the pears are densely clustered at the complex orchard. Therefore, the 3D stereo camera-based dataset combined with the Mask R-CNN vision algorithm had high accuracy in detecting the individual pears from gathered pears in a complex orchard environment.

Keywords: 3D stereo camera; Mask R-CNN; pear detection.

PubMed Disclaimer

Conflict of interest statement

The authors declare no conflict of interest.

Figures

**Figure 1**
Aerial view of orchards for data collection located at the Tsukuba-Plant Innovation Research Center (T-PIRC), University of Tsukuba, Tsukuba, Ibaraki. (a) satellite view of Tsukuba-Plant Innovation Research Center (T-PIRC); (b) the view of pear orchard in T-PIRC.

**Figure 2**
Different segmentation in pear detection using 3D camera datasets, (a) original image; (b) semantic segmentation; (c) object detection and (d) instance segmentation.

**Figure 3**
Mask R-CNN structure for pear quantity in orchards from 3D camera datasets. The original images enter the backbone network for selection and screening to get the feature maps. Then, the foreground and background are extracted in the RPN network, enter the ROI-Align network for standardization, and finally enter the head network to generate classes, boxes, and masks for pear detection.

**Figure 4**
The inner structure of ResNet101 as an example of second layers (C2): (a) is conv block and (b) is identity block. The images which were inputted into the ResNet have changed the channels. Conv block is the first stage of each layer, and the identity blocks and conv blocks were combined to the ResNet.

**Figure 5**
ResNet101 + FPN for pear quantity recognition.

**Figure 6**
RPN in Mask R-CNN for extracting proposals of original pear images.

**Figure 7**
Generation for IoU by comparing anchor boxes with ground truth boxes. If IoU > 0.7, then label = 1 positive; if IoU < 0.3, then label = −1 negative; others, label = 0.

**Figure 8**
Bilinear interpolation in ROI-Align was used to obtain fixed feature maps for pear recognition. P represents pixel coordinates that ROI-Align wanted to obtain after bilinear interpolation. Q11, Q12, Q22, and Q21 represent the four coordinates of known pixel points around point P.

**Figure 9**
Flow diagram of feature maps to produce boxes, classes, and masks for each pear in each fixed-size feature map after ROI-Align.

**Figure 10**
Pear prediction for determining FP, TN, TP, and FN using Mask R-CNN, (a) original image; (b) cv_mask input image before testing and (c) mask image after testing.

**Figure 11**
Mask R-CNN loss results from training losses and validation losses, (a) Total loss; (b) Mask R-CNN head bounding box loss; (c) Mask R-CNN head class loss; (d) Mask R-CNN mask loss; (e) RPN bounding box loss and (f) RPN class loss.

**Figure 12**
(a) Precision-recall curve of Faster R-CNN at learning rate = 0.001 in the testing set and (b) Precision-recall curve of Mask R-CNN at learning rate = 0.001 in the testing set.

**Figure 13**
Results of Mask-RCNN in different situations. Recognition of (a–c): separated pears in low light; (d–f): aggregated pears in low light; (g–i): separated pears in strong light, and (j–l) aggregated pears in strong light. (a,d,g,j) Original image; (b,e,h,k) Testing image in Mask R-CNN; and (c,f,i,l) Testing image in Faster R-CNN.

**Figure 14**
Results of Mask-R CNN in rotation angles. Recognition of (a–c): separated pear in low light; (d–f): aggregated pears in low light; (g–i): separated pears in strong light, and (j–l) aggregated pears in strong light (a,d,g,j) Original image; (b,e,h,k) Testing image in Mask R-CNN; and (c,f,i,l) Testing image in Faster R-CNN.

See this image and copyright information in PMC

Cited by

GreenFruitDetector: Lightweight green fruit detector in orchard environment.
Wang J, Shang Y, Zheng X, Zhou P, Li S, Wang H. Wang J, et al. PLoS One. 2024 Nov 14;19(11):e0312164. doi: 10.1371/journal.pone.0312164. eCollection 2024. PLoS One. 2024. PMID: 39541312 Free PMC article.
A novel hand-eye calibration method of picking robot based on TOF camera.
Zhang X, Yao M, Cheng Q, Liang G, Fan F. Zhang X, et al. Front Plant Sci. 2023 Jan 17;13:1099033. doi: 10.3389/fpls.2022.1099033. eCollection 2022. Front Plant Sci. 2023. PMID: 36733593 Free PMC article.
High-precision object detection network for automate pear picking.
Zhao P, Zhou W, Na L. Zhao P, et al. Sci Rep. 2024 Jun 28;14(1):14965. doi: 10.1038/s41598-024-65750-6. Sci Rep. 2024. PMID: 38942940 Free PMC article.
Intrarow Uncut Weed Detection Using You-Only-Look-Once Instance Segmentation for Orchard Plantations.
Sampurno RM, Liu Z, Abeyrathna RMRD, Ahamed T. Sampurno RM, et al. Sensors (Basel). 2024 Jan 30;24(3):893. doi: 10.3390/s24030893. Sensors (Basel). 2024. PMID: 38339611 Free PMC article.
Development of a Deep Learning Model for the Analysis of Dorsal Root Ganglion Chromatolysis in Rat Spinal Stenosis.
Li M, Zheng H, Koh JC, Choe GY, Choi EJ, Nahm FS, Lee PB. Li M, et al. J Pain Res. 2024 Apr 6;17:1369-1380. doi: 10.2147/JPR.S444055. eCollection 2024. J Pain Res. 2024. PMID: 38600989 Free PMC article.

References

1. Barua S. Understanding Coronanomics: The Economic Implications of the Coronavirus (COVID-19) Pandemic. [(accessed on 1 April 2020)]. Available online: https://ssrn.com/abstract=3566477.
1. Saito T. Advances in Japanese pear breeding in Japan. Breed. Sci. 2016;66:46–59. doi: 10.1270/jsbbs.66.46. - DOI - PMC - PubMed
1. Schrder C. Employment in European Agriculture: Labour Costs, Flexibility and Contractual Aspects. 2014. [(accessed on 1 April 2020)]. Available online: agricultura.gencat.cat/web/.content/de_departament/de02_estadistiques_ob....
1. Wei X., Jia K., Lan J., Li Y., Zeng Y., Wang C. Automatic method of fruit object extraction under complex agricultural background for vision system of fruit picking robot. Optik. 2004;125:5684–5689. doi: 10.1016/j.ijleo.2014.07.001. - DOI
1. Bechar A., Vigneault C. Agricultural robots for field operations: Concepts and components. Biosyst. Eng. 2016;149:94–111. doi: 10.1016/j.biosystemseng.2016.06.014. - DOI

MeSH terms

Actions
Actions
Actions
Actions

Grants and funding

21K05844/Japan Society for the Promotion of Science

LinkOut - more resources

Full Text Sources

Save citation to file

Email citation

Add to Collections

Add to My Bibliography

Your saved search

Create a file for external citation management software

Your RSS Feed

Pear Recognition in an Orchard from 3D Stereo Camera Datasets to Develop a Fruit Picking Mechanism Using Mask R-CNN

Affiliations

Pear Recognition in an Orchard from 3D Stereo Camera Datasets to Develop a Fruit Picking Mechanism Using Mask R-CNN

Authors

Affiliations

Abstract

Conflict of interest statement

Figures

Similar articles

Cited by

References

MeSH terms

Grants and funding

LinkOut - more resources

Full Text Sources