| ---
|
| license: apache-2.0
|
| pipeline_tag: image-segmentation
|
| tags:
|
| - infrared-small-target-detection
|
| - multimodal
|
| - vision-language
|
| - pytorch
|
| ---
|
|
|
|
|
| <h1>ADGNet: Asymmetric Dual-text Guided Network for Infrared Small Target Detection</h1>
|
|
|
| <p>
|
| <b>Tongtong Wang</b><sup>1</sup>
|
| <b>Mingzhu Xu</b><sup>1β</sup>
|
| <b>Chenglong Yu</b><sup>1</sup>
|
| <b>Jing Wang</b><sup>1</sup>
|
| <b>Xiaohui Lin</b><sup>1</sup>
|
| <b>Weili Guan</b><sup>2</sup>
|
| </p>
|
|
|
| <p>
|
| <sup>1</sup>School of Software, Shandong University<br>
|
| <sup>2</sup>Harbin Institute of Technology, Shenzhen<br>
|
| <sup>β</sup>Corresponding author
|
| </p>
|
|
|
| <p>
|
| <a href="<paper-link>">
|
| <img src="https://img.shields.io/badge/ACM%20MM-2026-blue" alt="ACM MM 2026">
|
| </a>
|
| <a href="https://github.com/iLearn-Lab/MM26-ADGNet">
|
| <img src="https://img.shields.io/badge/GitHub-MM26--ADGNet-black?logo=github" alt="GitHub">
|
| </a>
|
| </p>
|
|
|
| ## π Model Description
|
|
|
| This repository provides the official model checkpoints for **ADGNet: Asymmetric Dual-text Guided Network for Infrared Small Target Detection**, accepted by **ACM Multimedia 2026**.
|
|
|
| Infrared Small Target Detection aims to accurately segment weak and tiny targets from complex infrared backgrounds. Existing pure-vision methods rely mainly on pixel-level information, while existing vision-language methods commonly describe targets and backgrounds using a single textual prompt. Such a symmetric design overlooks the inherent semantic differences between sparse infrared targets and structurally complex backgrounds.
|
|
|
| ADGNet addresses this problem through three main components:
|
|
|
| - **Asymmetric Dual-text Prompt (ADP):** uses an abstract, image-independent target prompt and a detailed, image-dependent background prompt.
|
| - **Asymmetric Dual-Branch Interaction (ADBI):** independently performs target localization and background suppression using their corresponding textual priors.
|
| - **Adaptive Feature Aggregation (AFA):** dynamically fuses target-enhanced and background-suppressed features for accurate segmentation.
|
|
|
| The model uses the pretrained **CLIP ViT-B/16** text encoder to extract semantic representations from the target and background prompts.
|
|
|
| ---
|
|
|
| ## π Available Checkpoints
|
|
|
| All ADGNet checkpoints are hosted in this Hugging Face model repository.
|
|
|
| Download the required checkpoint directly from the **Files and versions** section of this repository.
|
|
|
| | Dataset | Checkpoint |
|
| | :--------: | :----------------------------------------------------------: |
|
| | IRSTD-1K | [`ADGNet_mIoU_72.38_IRSTD-1K.pth.tar`](https://huggingface.co/iLearn-Lab/MM26-ADGNet/blob/main/ADGNet_mIoU_72.38_IRSTD-1K.pth.tar) |
|
| | NUDT-SIRST | [`ADGNet_mIoU_95.53_NUDT-SIRST.pth.tar`](https://huggingface.co/iLearn-Lab/MM26-ADGNet/blob/main/ADGNet_mIoU_95.53_NUDT-SIRST.pth.tar) |
|
| | SIRST | [`ADGNet_mIoU_83.08_SIRST.pth.tar`](https://huggingface.co/iLearn-Lab/MM26-ADGNet/blob/main/ADGNet_mIoU_83.08_SIRST.pth.tar) |
|
|
|
| ---
|
|
|
| ## π Usage
|
|
|
| These checkpoints are designed to be used with the official ADGNet implementation:
|
|
|
| ```text
|
| https://github.com/iLearn-Lab/MM26-ADGNet
|
| ```
|
|
|
| ### 1. Clone the Official Repository
|
|
|
| ```bash
|
| git clone https://github.com/iLearn-Lab/MM26-ADGNet.git
|
| cd MM26-ADGNet
|
| ```
|
|
|
| ### 2. Prepare the Checkpoints
|
|
|
| Place the downloaded checkpoints in:
|
|
|
| ```text
|
| MM26-ADGNet/
|
| βββ SOTA_pth/
|
| βββ ADGNet_mIoU_72.38_IRSTD-1K.pth.tar
|
| βββ ADGNet_mIoU_95.53_NUDT-SIRST.pth.tar
|
| βββ ADGNet_mIoU_83.08_SIRST.pth.tar
|
| ```
|
|
|
| ### 3. Prepare the CLIP Text Encoder
|
|
|
| ADGNet uses the pretrained **CLIP ViT-B/16** model:
|
|
|
| ```bash
|
| git clone https://huggingface.co/openai/clip-vit-base-patch16
|
| ```
|
|
|
| Update the local CLIP model path in the corresponding project configuration or source file before inference.
|
|
|
| ### 4. Run Evaluation
|
|
|
| Example evaluation on IRSTD-1K:
|
|
|
| ```bash
|
| python train.py \
|
| --trainset "IRSTD-1K" \
|
| --testset "IRSTD-1K" \
|
| --dataset_dir "./datasets" \
|
| --mode test \
|
| --ckpt "./SOTA_pth/ADGNet_mIoU_72.38_IRSTD-1K.pth.tar"
|
| ```
|
|
|
| Replace the dataset name and checkpoint path when evaluating on NUDT-SIRST or SIRST.
|
|
|
| ---
|
|
|
| ## π Dataset and Text Annotation Preparation
|
|
|
| The original infrared images and ground-truth masks are not included in this model repository. Please obtain **IRSTD-1K**, **NUDT-SIRST**, and **SIRST** from their respective official sources.
|
|
|
| The asymmetric text annotations used by ADGNet are released separately in our Hugging Face dataset repository:
|
|
|
| - **AITIR Text Annotations:** [`Download`](<huggingface-text-dataset-link>)
|
|
|
| After downloading the original datasets and text annotations, organize them according to the official ADGNet repository:
|
|
|
| ```text
|
| datasets/
|
| βββ IRSTD-1K/
|
| β βββ images/
|
| β βββ masks/
|
| β βββ img_idx/
|
| β βββ text/
|
| βββ NUDT-SIRST/
|
| β βββ images/
|
| β βββ masks/
|
| β βββ img_idx/
|
| β βββ text/
|
| βββ SIRST/
|
| βββ images/
|
| βββ masks/
|
| βββ img_idx/
|
| βββ text/
|
| ```
|
|
|
| ---
|
|
|
| ## π― Intended Use
|
|
|
| The released checkpoints are intended for:
|
|
|
| - Academic research on infrared small target detection
|
| - Reproduction of the results reported in the ADGNet paper
|
| - Evaluation on IRSTD-1K, NUDT-SIRST, and SIRST
|
| - Research on multimodal and text-guided infrared image segmentation
|
| - Comparison with other infrared small target detection methods
|
|
|
| ---
|
|
|
| ## β οΈ Limitations
|
|
|
| - The model requires both infrared images and corresponding textual prompts.
|
| - Detection performance may vary when applied to datasets or scenes that differ substantially from the training distribution.
|
| - The released checkpoints are designed for the dataset splits and evaluation settings used in the paper.
|
| - The model depends on the pretrained CLIP ViT-B/16 text encoder.
|
| - The original infrared datasets are subject to their respective licenses and terms of use.
|
|
|
| ---
|
|
|
| ## π Related Resources
|
|
|
| - **GitHub Repository:** [iLearn-Lab/MM26-ADGNet](https://github.com/iLearn-Lab/MM26-ADGNet)
|
| - **Paper:** [`ADGNet: Asymmetric Dual-text Guided Network for Infrared Small Target Detection`](<paper-link>)
|
| - **Text Annotations:** [`AITIR Text Annotations`](<huggingface-text-dataset-link>)
|
|
|
| ---
|
|
|
| ## π Citation
|
|
|
| If you find ADGNet or the released checkpoints useful in your research, please consider citing our paper:
|
|
|
| ```bibtex
|
|
|
| ```
|
|
|
| Please also consider checking out and citing our other related work:
|
|
|
| ```bibtex
|
|
|
| ```
|
|
|
| ---
|
|
|