Abstract
Few-shot multispectral object detection (FSMOD) addresses the challenge ofdetecting objects across visible and thermal modalities with minimal annotateddata. In this paper, we explore this complex task and introduce a frameworknamed "FSMODNet" that leverages cross-modality feature integration to improvedetection performance even with limited labels. By effectively combining theunique strengths of visible and thermal imagery using deformable attention, theproposed method demonstrates robust adaptability in complex illumination andenvironmental conditions. Experimental results on two public datasets showeffective object detection performance in challenging low-data regimes,outperforming several baselines we established from state-of-the-art models.All code, models, and experimental data splits can be found athttps://anonymous.4open.science/r/Test-B48D.