CMCSENet

Published: 8 July 2026| Version 1 | DOI: 10.17632/v2cmt87777.1
Contributor:
Guoquan Guo

Description

RGB-IR object detection effectively improves the robustness of object detection in complex scenarios by fusing the texture information of visible images and the thermal radiation information of infrared images. However, existing multimodal detection methods still face two key challenges. On the one hand, single-modality features and fused features often contain target responses, background interference, and local noise simultaneously. Directly extracting multi-scale contextual information from such entangled features is not only inefficient, but also prone to introducing noise interference. On the other hand, cross modal feature fusion is easily affected by modality discrepancies, pseudo thermal sources, and texture blur, which weakens the reliability and discriminability of fused semantics. To address these issues, we proposes a Cross Modal Context and Semantic Enhancement Network (CMCSENet). Specifically, we design a Region Guided Context Modeling Block (RGCMB), which organizes complex features into clearer candidate region representations through region guided mapping, and combines main context kernels with sparse context sub-kernels in the context modeling branches to simultaneously enhance local target details, regional structures, and large-range semantic consistency. Furthermore, we propose a Cross Modal Semantic Enhancement Module (CSEM), which progressively calibrates fused features from four aspects, namely modality response preservation, statistical semantic modeling, spatial structure enhancement, and frequency discriminative screening. In this way, cross modal noise and irrelevant background responses are suppressed, and the stability and discriminability of fused features are improved.

Files

Categories

Object Detection, Feature Fusion

Licence