Vision Transformers for Satellite Images, Without the Jargon

· 3 min read · Syed Omar Faruk Towaha
Vision Transformers for Satellite Images, Without the Jargon

Your phone camera sees three colours: red, green and blue. Many Earth-observation satellites see several more, including light your eyes can't see at all. Plants, for example, reflect a lot of near-infrared light, so in the right band a healthy field glows like a neon sign.

That's one reason satellite images are such a good playground for machine learning, and one reason they're trickier than holiday photos.

Spectral bands
Approximate band centres for a typical multispectral sensor.

What we want from these images

Questions like:

Each of these can be framed as segmentation: label every pixel (or small area) as water, vegetation, built-up land and so on.

Enter the vision transformer

Transformers were originally built for text: they read a sentence as a sequence of tokens and let every token "pay attention" to every other token. A vision transformer does the same with images. It cuts the image into small square patches, turns each patch into a token, and lets the patches compare notes.

How a ViT reads
Patches become tokens; attention lets distant patches inform each other.

Why is that helpful for satellite data? Because context matters. A dark patch could be water, shadow or fresh asphalt. Looking at the surrounding patches (a river shape, a city grid, a mountain) helps decide.

Why satellite data is weird

Mixing physics with deep learning

One approach I find promising is combining physical knowledge with learned models. Scientists have long used spectral indices such as NDVI, which compares near-infrared and red light to measure vegetation:

import numpy as np

def ndvi(nir: np.ndarray, red: np.ndarray) -> np.ndarray:
    return (nir - red) / (nir + red + 1e-6)

Feeding the model physically meaningful inputs like these, alongside raw bands, can help it learn faster from fewer labels and behave more sensibly on new regions. Physics provides the common sense; the transformer provides the pattern recognition.

Why it matters

Satellites revisit the same places every few days, for free or nearly free in some programmes. With good models, that becomes a continuous health check for lakes, coastlines, crops and forests. For countries facing floods and climate change, including my home country, that's not a research curiosity. It's an early-warning system.

// related

// prefer the terminal?

Open the terminal blog and type read vision-transformers-satellite-images-no-jargon.