P2Fusion: Prompt-based Progressive Infrared-Visible Image Fusion via Dual-Prior Distillation

Yi Shi1,*, Huichao Xie1,*, Yuqing Wang1, Mingyu Wang1, Kaihui Yang1, Yu Liu2, Ruitao Lu3, Lizhe Li1, Junwei Han1,4,†, Dingwen Zhang1,†
*These authors contributed equally   Corresponding authors
1Northwestern Polytechnical University,
2Hefei University of Technology,
3Rocket Force University of Engineering,
4Chongqing University of Posts and Telecommunications,
P2Fusion Teaser
Motivation and Performance Overview. (a) Prior integration paradigm (Left): Existing methods inject priors as static constraints, often causing optimization conflicts. Our P2Fusion provides dynamic, prompt-based guidance via image-intrinsic priors to mitigate gradient interference. (b) Fusion quality (Middle): Radar charts across four datasets (MSRS, M3FD, FMB, RoadScene) demonstrate our superior visual fidelity. (c) Downstream perception (Right): P2Fusion consistently enhances object detection and semantic segmentation performance on FMB and M3FD benchmarks.

Abstract

Infrared-visible image fusion (IVIF) is pivotal for multimodal perception, yet reconciling the inherent information disparity between thermal and textual features remains a fundamental challenge. Existing prior-guided methods often rely on static constraints that induce optimization conflicts or utilize extrinsic semantic priors from large-scale foundation models (e.g., CLIP/DINO), which frequently fail to exploit the intrinsic modality characteristics essential for high-fidelity fusion. To address these issues, we propose P2Fusion, a prior-guided distillation-based framework that reformulates IVIF via dual intrinsic prompts. Instead of imposing hard-coded penalties, we distill image-intrinsic priors, thermal saliency and spatial quality into learnable, dynamic regulators. Specifically, a Teach-to-Fuse mechanism provides dual-granularity progressive guidance, coupled with a Gated Dynamic Expert Recalibration (GDER) module for decoupled feature refinement. This design enables the network to adaptively mediate modal competition through expert specialization. Extensive experiments demonstrate that P2Fusion achieves state-of-the-art performance across five mainstream datasets. Notably, our framework demonstrates consistent performance advantages in fusion quality, achieving state-of-the-art results in 14 out of 20 key evaluation metrics across 5 benchmarks. Furthermore, it effectively contributes to the robustness of downstream perception, such as +3.2% mAP on MSRS, +0.5% mAP on M3FD and +0.9% mAP on DroneVehicle for object detection.

Overview

Overview of the P2Fusion framework
Illustration of the proposed framework. (a) Overall pipeline: prior-guided visible and infrared features are fused via cross-attention, where the proposed GDER module adaptively refines modality-specific and interaction features for reconstruction. (b) Spatial Quality Distillation: a quality assessor provides patch-level quality guidance to supervise spatial-detail learning. (c) Thermal Saliency Distillation: a saliency teacher offers thermal saliency priors to enhance target-aware fusion.

Main Results

Main results of P2Fusion
Main results of P2Fusion
Main results of P2Fusion

Visualization

Qualitative comparison across diverse scenes
Qualitative comparison between our P2Fusion and existing image fusion methods. From top to bottom: well-lit scenes from M3FD, smoke-occluded scenes from FMB, and over-exposed scenes from RoadScene.
Detection and segmentation qualitative comparison
Qualitative comparison of P2Fusion against other fusion methods on the downstream object detection task (M3FD) and the semantic segmentation task (FMB); P2Fusion consistently demonstrates superior performance.

3D Model