Instructions to use raja0007/satquery-satellite-vlm with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use raja0007/satquery-satellite-vlm with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("OpenGVLab/InternVL3_5-1B-Instruct") model = PeftModel.from_pretrained(base_model, "raja0007/satquery-satellite-vlm") - Notebooks
- Google Colab
- Kaggle
๐ฐ๏ธ SatQuery AI โ Multimodal Satellite Vision-Language Model
SatQuery is a parameter-efficient fine-tuned Vision-Language Model (VLM) engineered for multi-sensor Earth Observation, disaster response, and change detection.
Fine-tuned on top of OpenGVLab/InternVL3_5-1B-Instruct across 16,000 multimodal pairs (32,000 Sentinel-1 SAR and Sentinel-2 Optical images, spanning 80,000 VQA instances), SatQuery simultaneously understands:
- ๐ก Sentinel-1 SAR (Synthetic Aperture Radar): Co-pol (VV), cross-pol (VH), polarimetric difference, surface roughness, and water specular reflection.
- ๐ท Sentinel-2 Multispectral & Optical: True color RGB, Shortwave Infrared (SWIR), False Color NIR, and Scene Classification.
๐ Benchmark Evaluation Results
Evaluated on the official held-out test split of 2,000 multi-sensor image pairs across 3 operational sensor modes:
| Metric / Category | Score |
|---|---|
| Soft Token F1 Score | 56.19% (Project Record) |
| Exact Match (EM) | 48.00% (24 / 50) |
| Land Cover Classification | 83.33% (5 / 6) |
| Feature Counting | 80.00% (4 / 5) |
| Presence / Absence Verification | 80.00% (4 / 5) |
| Area Coverage Estimation | 66.67% (4 / 6) |
| Bounding Box Syntax Valid Rate | 100.00% |
| Repetition Loop Rate | 0.00% |
| EOS Termination Rate | 100.00% |
Sensor Modality Ablation:
- Joint Multimodal (Sentinel-1 SAR + Sentinel-2 Optical): 56.19% Soft F1
- Radar-Only Mode (Sentinel-1 SAR): 58.42% Soft F1
- Optical-Only Mode (Sentinel-2 Optical): 56.44% Soft F1
โก Instant Ready-to-Run Usage (Using the Uploaded Code)
You don't need to write model loading code! Simply download inference.py directly from this repository:
# Download the standalone inference engine
curl -O https://huggingface.co/raja0007/satquery-satellite-vlm/raw/main/inference.py
# Query any optical satellite image:
python inference.py --image satellite_rgb.png --modality optical --question "What land cover is visible?"
# Query SAR radar (Sentinel-1):
python inference.py --image radar.png --modality sar --question "Is active flood inundation present?"
# Multimodal SAR + Optical analysis:
python inference.py --sar sar.png --optical optical.png --question "Is there flooding along the river?"
# Interactive chat loop:
python inference.py --interactive
Or Use the Python Class:
from inference import SatQueryModel
model = SatQueryModel()
answer = model.query(image="flood_scene.png", question="Is there active water inundation?", modality="optical")
print("SatQuery:", answer)
๐ Advanced Custom Pipeline (Inference in Python)
1. Installation
pip install torch torchvision transformers peft bitsandbytes tifffile pillow
2. Loading the Model Manually (LoRA Adapter)
import torch
from transformers import AutoModel, AutoTokenizer, BitsAndBytesConfig
from peft import PeftModel
from torchvision import transforms
from PIL import Image
# 1. Load Base Architecture (4-bit NF4 quantized for low VRAM)
base_repo = "OpenGVLab/InternVL3_5-1B-Instruct"
adapter_repo = "raja0007/satquery-satellite-vlm"
bnb_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.bfloat16,
bnb_4bit_use_double_quant=True,
)
tokenizer = AutoTokenizer.from_pretrained(base_repo, trust_remote_code=True)
base_model = AutoModel.from_pretrained(
base_repo,
quantization_config=bnb_config,
torch_dtype=torch.bfloat16,
trust_remote_code=True,
device_map="auto",
)
# Connect fine-tuned satellite adapter
model = PeftModel.from_pretrained(base_model, adapter_repo)
model.eval()
# Vision dtype
vision_dtype = next(model.base_model.model.vision_model.parameters()).dtype
3. Running Multi-Sensor Satellite Inference
# Transform for input images
transform = transforms.Compose([
transforms.Resize((448, 448)),
transforms.ToTensor(),
transforms.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]),
])
# Load your Sentinel-1 SAR and Sentinel-2 Optical images (RGB PIL Images)
sar_image = Image.open("path_to_sar_image.png").convert("RGB")
opt_image = Image.open("path_to_optical_image.png").convert("RGB")
t_s1 = transform(sar_image).unsqueeze(0).to(model.device, dtype=vision_dtype)
t_s2 = transform(opt_image).unsqueeze(0).to(model.device, dtype=vision_dtype)
pixel_values = torch.cat([t_s1, t_s2], dim=0)
question = "Is there floodwater or inundation present in this scene?"
# Structured multimodal prompt
prompt = (
f"<MODALITY: Sentinel-1 SAR Radar>\n<image>\n"
f"<MODALITY: Sentinel-2 Optical>\n<image>\n"
f"Question: {question}\nAnswer:"
)
with torch.no_grad():
response = model.chat(
tokenizer=tokenizer,
pixel_values=pixel_values,
question=prompt,
num_patches_list=[1, 1],
generation_config=dict(max_new_tokens=64, do_sample=False, temperature=0.0),
)
print("SatQuery Answer:", response)
๐ ๏ธ Training Specifications
- Base Model:
OpenGVLab/InternVL3_5-1B-Instruct - LoRA Target Modules:
q_proj,k_proj,v_proj,o_proj,gate_proj,up_proj,down_proj - LoRA Parameters: Rank $r=16$, $\alpha=32$, Dropout $0.05$ (Trainable parameters: ~38 MB)
- Visual Token Budget: 512 tokens ($256 \text{ tokens (SAR)} + 256 \text{ tokens (Optical)}$)
- Global Optimization Steps: 1,000 steps with effective batch size 16 (16,000 pairs)
- Training Time: 7.23 hours on NVIDIA GeForce RTX 5060 Laptop GPU (8 GB VRAM)
- Final Training Loss: 0.2347 | Best Validation Loss: 0.2782
๐ Citation & Credits
Fine-tuned by Raja as part of the SatQuery AI initiative for multi-modal Earth Observation reasoning.
Based on the InternVL-3.5 architecture by OpenGVLab.
- Downloads last month
- 40
Model tree for raja0007/satquery-satellite-vlm
Base model
OpenGVLab/InternVL3_5-1B-PretrainedDataset used to train raja0007/satquery-satellite-vlm
Evaluation results
- Soft Token F1 on BigEarthNet Multimodal Held-Out Test Splitself-reported56.190
- Exact Match (EM) on BigEarthNet Multimodal Held-Out Test Splitself-reported48.000