BIO597 · Spatial Analysis of Biodiversity

Species distribution models
from first principles

A conceptual bridge from occurrence records and environmental rasters to mapped habitat suitability

Lecture 06 · Fall 2026
Today’s questions
  1. What is the basic logic behind a species distribution model?
  2. How do occurrence records become information about environmental conditions?
  3. How can a suitability rule be projected back onto a map?
  4. Why is suitability not the same thing as probability of occurrence?
Motivation

We often want to map places we have not sampled

Occurrence records are uneven glimpses of where a species has been observed.

An SDM asks whether those observations are associated with environmental conditions that occur elsewhere.

The SDM idea

Connect observations to environments, then project that relationship

OccurrencesCoordinates where the species was recorded
→
EnvironmentRaster values at those locations
→
Suitability mapScores for other cells with similar conditions

The lab keeps this chain visible by using one species, one predictor, and one equation.

Important distinction

Observed does not mean all suitable places are known

Presence records

  • Species was recorded there
  • Usually affected by sampling effort
  • May miss many suitable locations

Absence records

  • Species was searched for and not detected
  • Requires survey design or detection assumptions
  • Not the same as a blank spot on a map
Background

Background is the environment available in the study area

A background sample is not a set of absences.

It tells us what environmental values exist across the region where we are making predictions.

  • Occurrences: used environments
  • Background: available environments
  • SDM signal: used differs from available
One predictor

Start with annual mean temperature

SimpleOne axis is easy to inspect
SpatialTemperature varies across Maine
BiologicalTemperature can affect activity and survival
IncompleteMany important processes are missing

A deliberately simple model is useful because every assumption is visible.

Environmental space

Compare occurrence values with background values

Cooler Annual mean temperature Warmer

Teal bars represent available background; coral overlays represent occurrence records.

A first-principles rule

Suitability is highest near the observed mean

Today’s model:

Places are more suitable when annual mean temperature is close to the mean annual mean temperature at occurrence records.

Annual mean temperature Suitability
The equation

A simple curve makes the assumption explicit

def temperature_suitability(temperature, optimum, tolerance):
    z = (temperature - optimum) / tolerance
    return np.exp(-0.5 * z**2)

optimum = occurrence_temperature.mean()
tolerance = occurrence_temperature.std()

suitability = temperature_suitability(raster_temperature, optimum, tolerance)

This is not the only possible curve. It is a transparent starting point.

Tolerance

One parameter changes the breadth of prediction

Narrow
Observed SD
Broad

A broad tolerance predicts more suitable area; a narrow tolerance predicts less.

Projection

Every raster cell gets passed through the same rule

Cell temperatureExample: 6.4 C
→
Suitability functionCompare to optimum and tolerance
→
Cell scoreExample: 0.82

Projection is just repeated prediction across the grid.

Thresholding

A binary map hides a modeling choice

Continuous suitability

  • Retains gradient information
  • Shows relative differences among cells
  • Often better for interpretation

Suitable / unsuitable

  • Easy to communicate
  • Depends on an arbitrary or justified cutoff
  • Can exaggerate certainty
Interpretation

Suitability is not probability of presence

Suitability

A relative score based on similarity to modeled conditions.

Occurrence

Depends on environment, dispersal, history, detection, and sampling effort.

Probability

Requires stronger statistical assumptions and usually validation data.

A high score means "similar to known occurrence environments" in this lab.

What we are leaving out

A one-variable model is a sketch

Sampling biasRecords cluster near roads, people, and projects
DetectionNot seeing a species is not always absence
ProcessDispersal, interactions, habitat, and history matter
ScaleRaster cells miss fine-scale microhabitats
What SDM packages add

The same idea with more machinery

  • Multiple predictors and interactions
  • Flexible response curves
  • Different sampling designs and background strategies
  • Cross-validation and evaluation metrics
  • Uncertainty, model comparison, and reproducible prediction workflows

The lab gives us the mental model before the machinery arrives.

Takeaways for the lab exercise

Today we build the smallest useful SDM

  1. Extract one environmental value at occurrence records.
  2. Compare occurrence environments with background environments.
  3. Define a transparent suitability curve.
  4. Project that curve back across the raster.
  5. Interpret the map cautiously.

Next: open the lab and make the SDM logic visible in code.