A multimodal framework for pepper diseases and pests detection

一种用于辣椒病虫害检测的多模态框架

阅读:2

Abstract

Pepper diseases and pests typically exhibit small target proportions, diverse shapes and sizes, complex imaging backgrounds, and similarities with the background. Existing detection methods perform poorly in identifying targets of different sizes and shapes within the same scene, and they lack adequate noise suppression capabilities. To address the practical needs of detecting pepper diseases and pests in complex scenarios, we have constructed the first multimodal pepper diseases and pests object detection dataset (PDD). This dataset includes a wide variety of diseases and pests images, along with detailed natural language descriptions of their attributes. Locating the described targets in complex scenes with similar disease symptoms and leaf occlusion presents a significant challenge. To tackle this issue, we propose the PepperNet model for object detection in pepper diseases and pests images using natural language descriptions. This model decomposes complex multimodal features of language and images into explicit attribute features and employs fine-grained multimodal attribute contrast learning strategies. This approach effectively distinguishes subtle local differences between similar objects, achieving fine-grained mapping from language to vision in complex scenarios. Our detection results show a mAP@0.5 of 91.93% and a detection speed of 121.8 frames per second. Visualizations indicate that the model maintains high robustness under varying noise levels and occlusion conditions, demonstrating superior performance and stability across diverse complex scenarios.

特别声明

1、本页面内容包含部分的内容是基于公开信息的合理引用;引用内容仅为补充信息,不代表本站立场。

2、若认为本页面引用内容涉及侵权,请及时与本站联系,我们将第一时间处理。

3、其他媒体/个人如需使用本页面原创内容,需注明“来源:[生知库]”并获得授权;使用引用内容的,需自行联系原作者获得许可。

4、投稿及合作请联系:info@biocloudy.com。