๐Ÿ›ฐ๏ธ SatQuery AI โ€” Multimodal Satellite Vision-Language Model

SatQuery is a parameter-efficient fine-tuned Vision-Language Model (VLM) engineered for multi-sensor Earth Observation, disaster response, and change detection.

Fine-tuned on top of OpenGVLab/InternVL3_5-1B-Instruct across 16,000 multimodal pairs (32,000 Sentinel-1 SAR and Sentinel-2 Optical images, spanning 80,000 VQA instances), SatQuery simultaneously understands:

  1. ๐Ÿ“ก Sentinel-1 SAR (Synthetic Aperture Radar): Co-pol (VV), cross-pol (VH), polarimetric difference, surface roughness, and water specular reflection.
  2. ๐Ÿ“ท Sentinel-2 Multispectral & Optical: True color RGB, Shortwave Infrared (SWIR), False Color NIR, and Scene Classification.

๐Ÿ“Š Benchmark Evaluation Results

Evaluated on the official held-out test split of 2,000 multi-sensor image pairs across 3 operational sensor modes:

Metric / Category Score
Soft Token F1 Score 56.19% (Project Record)
Exact Match (EM) 48.00% (24 / 50)
Land Cover Classification 83.33% (5 / 6)
Feature Counting 80.00% (4 / 5)
Presence / Absence Verification 80.00% (4 / 5)
Area Coverage Estimation 66.67% (4 / 6)
Bounding Box Syntax Valid Rate 100.00%
Repetition Loop Rate 0.00%
EOS Termination Rate 100.00%

Sensor Modality Ablation:

  • Joint Multimodal (Sentinel-1 SAR + Sentinel-2 Optical): 56.19% Soft F1
  • Radar-Only Mode (Sentinel-1 SAR): 58.42% Soft F1
  • Optical-Only Mode (Sentinel-2 Optical): 56.44% Soft F1

โšก Instant Ready-to-Run Usage (Using the Uploaded Code)

You don't need to write model loading code! Simply download inference.py directly from this repository:

# Download the standalone inference engine
curl -O https://huggingface.co/raja0007/satquery-satellite-vlm/raw/main/inference.py

# Query any optical satellite image:
python inference.py --image satellite_rgb.png --modality optical --question "What land cover is visible?"

# Query SAR radar (Sentinel-1):
python inference.py --image radar.png --modality sar --question "Is active flood inundation present?"

# Multimodal SAR + Optical analysis:
python inference.py --sar sar.png --optical optical.png --question "Is there flooding along the river?"

# Interactive chat loop:
python inference.py --interactive

Or Use the Python Class:

from inference import SatQueryModel

model = SatQueryModel()
answer = model.query(image="flood_scene.png", question="Is there active water inundation?", modality="optical")
print("SatQuery:", answer)

๐Ÿš€ Advanced Custom Pipeline (Inference in Python)

1. Installation

pip install torch torchvision transformers peft bitsandbytes tifffile pillow

2. Loading the Model Manually (LoRA Adapter)

import torch
from transformers import AutoModel, AutoTokenizer, BitsAndBytesConfig
from peft import PeftModel
from torchvision import transforms
from PIL import Image

# 1. Load Base Architecture (4-bit NF4 quantized for low VRAM)
base_repo = "OpenGVLab/InternVL3_5-1B-Instruct"
adapter_repo = "raja0007/satquery-satellite-vlm"

bnb_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_compute_dtype=torch.bfloat16,
    bnb_4bit_use_double_quant=True,
)

tokenizer = AutoTokenizer.from_pretrained(base_repo, trust_remote_code=True)
base_model = AutoModel.from_pretrained(
    base_repo,
    quantization_config=bnb_config,
    torch_dtype=torch.bfloat16,
    trust_remote_code=True,
    device_map="auto",
)

# Connect fine-tuned satellite adapter
model = PeftModel.from_pretrained(base_model, adapter_repo)
model.eval()

# Vision dtype
vision_dtype = next(model.base_model.model.vision_model.parameters()).dtype

3. Running Multi-Sensor Satellite Inference

# Transform for input images
transform = transforms.Compose([
    transforms.Resize((448, 448)),
    transforms.ToTensor(),
    transforms.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]),
])

# Load your Sentinel-1 SAR and Sentinel-2 Optical images (RGB PIL Images)
sar_image = Image.open("path_to_sar_image.png").convert("RGB")
opt_image = Image.open("path_to_optical_image.png").convert("RGB")

t_s1 = transform(sar_image).unsqueeze(0).to(model.device, dtype=vision_dtype)
t_s2 = transform(opt_image).unsqueeze(0).to(model.device, dtype=vision_dtype)
pixel_values = torch.cat([t_s1, t_s2], dim=0)

question = "Is there floodwater or inundation present in this scene?"

# Structured multimodal prompt
prompt = (
    f"<MODALITY: Sentinel-1 SAR Radar>\n<image>\n"
    f"<MODALITY: Sentinel-2 Optical>\n<image>\n"
    f"Question: {question}\nAnswer:"
)

with torch.no_grad():
    response = model.chat(
        tokenizer=tokenizer,
        pixel_values=pixel_values,
        question=prompt,
        num_patches_list=[1, 1],
        generation_config=dict(max_new_tokens=64, do_sample=False, temperature=0.0),
    )

print("SatQuery Answer:", response)

๐Ÿ› ๏ธ Training Specifications

  • Base Model: OpenGVLab/InternVL3_5-1B-Instruct
  • LoRA Target Modules: q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
  • LoRA Parameters: Rank $r=16$, $\alpha=32$, Dropout $0.05$ (Trainable parameters: ~38 MB)
  • Visual Token Budget: 512 tokens ($256 \text{ tokens (SAR)} + 256 \text{ tokens (Optical)}$)
  • Global Optimization Steps: 1,000 steps with effective batch size 16 (16,000 pairs)
  • Training Time: 7.23 hours on NVIDIA GeForce RTX 5060 Laptop GPU (8 GB VRAM)
  • Final Training Loss: 0.2347 | Best Validation Loss: 0.2782

๐Ÿ“œ Citation & Credits

Fine-tuned by Raja as part of the SatQuery AI initiative for multi-modal Earth Observation reasoning. Based on the InternVL-3.5 architecture by OpenGVLab.

Downloads last month
40
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for raja0007/satquery-satellite-vlm

Dataset used to train raja0007/satquery-satellite-vlm

Evaluation results

  • Soft Token F1 on BigEarthNet Multimodal Held-Out Test Split
    self-reported
    56.190
  • Exact Match (EM) on BigEarthNet Multimodal Held-Out Test Split
    self-reported
    48.000