DOI: **.****/s*****-***-****-*
Potential of Multivariate Quantitative Methods for
Delineation and Visualization of Ecoregions
WILLIAM W. HARGROVE* from the limitations of human subjectivity, making possible a
FORREST M. HOFFMAN new array of ecologically useful derivative products. A
Environmental Sciences Division red green blue visualization based on principal components
Computer Science and Math Division analysis of ecoregion centroids indicates with color the rel-
Oak Ridge National Laboratory ative combination of environmental conditions found within
P. O. Box 2008, M.S. 6407 each ecoregion. Multiple geographic areas can be classified
Oak Ridge, Tennessee 37831-6407 into a single common set of quantitative ecoregions to pro-
vide a basis for comparison, or maps of a single area through
ABSTRACT / Multivariate clustering based on fine spatial time can be classified to portray climatic or environmental
resolution maps of elevation, temperature, precipitation, soil changes geographically in terms of current conditions.
characteristics, and solar inputs has been used at several Quantified representativeness can characterize borders be-
specified levels of division to produce a spectrum of tween ecoregions as gradual, sharp, or of changing char-
quantitative ecoregion maps for the conterminous United acter along their length. Similarity of any ecoregion to all other
States. The coarse ecoregion divisions accurately capture ecoregions can be quantified and displayed as a repre-
intuitively-understood regional environmental differences, sentativeness map. The representativeness of an existing
whereas the finer divisions highlight local condition spatial array of sample locations or study sites can be
gradients, ecotones, and clines. Such statistically generated mapped relative to a set of quantitative ecoregions, sug-
ecoregions can be produced based on user-selected gesting locations for additional samples or sites. In addition,
continuous variables, allowing customized regions to be the shape of Hutchinsonian niches in environment space can
delineated for any specific problem. By creating an objective be defined if a multivariate range map of species occurrence
ecoregion classification, the ecoregion concept is removed is available.
Ecoregions are designed to help users visualize and using human expertise in a qualitative, weight-of-evi-
understand similarities across complex multivariate dence approach (McMahon and others 2001).
environmental factors by grouping areas into like cat- It is not our intention to add to this debate. These
egories. The basis for such groupings has many con- conceptual issues are well represented in this special
tentious conceptual underpinnings. Debate abounds as issue and elsewhere. Rather, we sidestep the question
to whether ecoregions should be specialized for a as to whether ecoregions are computable and demon-
particular use or general purpose, spatially contiguous strate the rami cations of quantitatively deriving eco-
versus disjunct, nestable versus nonhierarchical, and regions. We argue that the spate of ancillary products
whether ecoregions can be defensible as units of resulting from the quantitative treatment of ecoregions
management, legislation, or even ecological triage enhances and expands the utility of the ecoregion
(Omernik 1995, 2003; Overton and others 2002; concept and makes substantial new contributions to
Leathwick and others 2003). Supreme among these niche modeling, network and sample design, change
issues, however, is the question of whether ecoregions detection, and conservation.
can (or should) be delineated using quantitative sta- Regionalizations are models, whether quantitatively
tistical methods or whether they can only be drawn or qualitatively derived. Quantitative models are, how-
ever, more explicit, repeatable, transferable, and
defensible than subjective models based on human
KEY WORDS: Clustering; Climate change; Ecotone; Environmental expertise. This transferability and repeatability makes
envelope; Fences; Gradient; Network; Niche; Pre- quantitative models more objective than their qualita-
serve design; Range; Representativeness; Similarity;
tive counterparts. Human experts might be able to
Time series
rationally defend drawing a particular borderline be-
tween ecoregions, but they might be unable to eluci-
Published online February 25, 2005.
date the method used to place it at that precise
*Author to whom correspondence should be addressed; email:
location. Quantitative ecoregionalization techniques
***@****.***.****.***
2005 Springer Science+Business Media, Inc.
Environmental Management Vol. 34, Suppl. 1, pp. S39 S60
S40 W. W. Hargrove and F. M. Hoffman
are not perfectly objective, because they still require others (1999) recently revisited the Life Zone approach
subjective ecological expertise in the choice of data for global vegetation, finding it adequate for forests, but
layers to include and in the interpretation of the of limited utility for grasslands, shrublands, and non-
resulting ecoregions. Nevertheless, the pursuit of a vegetated lands.
fully describable quantitative method for ecoregional- Kohonen s (1982, 1995) Self-Organizing Maps
ization is a desirable goal, whether as a replacement or (SOMs) use a two-layer neural network to divide a
an augmentation to qualitative, expert-opinion-based multivariate map into ecoregions. During training,
approaches. neurons are reinforced within a gradually shrinking
The ability to explicitly select and control the input network neighborhood around the best predictors
variables allows quantitative regionalizations to be cus- (Hung 1993). Using the rst ve principal components
tomized for speci c uses. Using quantitative methods, of seasonal averages of climate data from 18 stations,
special-purpose ecoregions can be created speci cally Malmgren and Winter (1999) used a one-dimensional
for particular uses. Such customized regionalizations SOM to divide Puerto Rico into four natural climatic
can provide additional discrimination in places where zones. It might be dif cult, however, to transfer the
learning from one neural network into a form that
popular generalized ecoregion schemes might miss
particular features of importance. can be used by others.
In qualitative ecoregionalization, human experts
Quantitative Ecoregions for Single Species
might intentionally adjust the weighting of particular
Mapping the geographic range of a single species of
input layers in their mental model in particular spots
animal or plant is a special case of ecoregion delinea-
across the map, whereas quantitative methods usually
tion. Ecological niche theory is central to understanding
produce regionalizations in which all input variables
how environmental change affects species abundance
receive equal weighting. Such evenly weighted ecore-
patterns (Jackson and Overpeck 2000). Hutchinson
gion products provide, at the very least, an initial basis
(1957) conceived of the niche as a multidimensional
from which factor weights can be subsequently spatially
hypervolume with dimensions de ned by the envi-
modi ed by human expertise. Explicit spatial maps of
altered weights, if known, could be considered directly ronmental factors that in uence the tness of individ-
uals of that species.
as inputs into a fully quantitative model.
Although Hutchinson s conception was static, envi-
ronmental change involves temporal alteration of
Sampling the Toolbox of Quantitative Methods
combinations of niche variables. Hutchinson s niche
envelope assumes that a steady-state equilibrium with
A full review of quantitative methods that have been
present environmental conditions has allowed ade-
used in regionalization is beyond the scope of this
article. Although not exhaustive, brief treatment of quate time for perfect adaptation and exhaustive
major types of quantitative approach might provide migration to all parts of the potential range. Acclima-
tion to changing conditions and historical factors lim-
context, and counterbalance more typical subjective
iting geographic dispersal are not considered by most
approaches described elsewhere in this special issue.
quantitative range-prediction techniques [but migra-
Most quantitative approaches rely on Gleasonian rela-
tion can be simulated in a subsequent step (i.e., Pet-
tionships between environmental patterns and species
erson and others 2003)].
occurrence, but some have been tested directly using
Generalized Additive Modeling (GAM) uses regres-
species distribution data, which include Clementsian
effects of biotic interactions on geographic distribu- sion modeling to establish empirical relationships be-
tions. Strengths and weaknesses of each quantitative tween a response variable (i.e., presence of a particular
species at a location) and an individually smoothed set
method are highlighted.
of spatial predictor variables (Hastie and Tibshirani
The Holdridge Life Zone model, usually shown as a
1990). GAM additively calculates the component re-
set of hexagons arranged inside a triangular plot, was
sponse and can handle nonlinear and nonmonotonic
an early quantitative multivariate approach for de ning
relationships between the response and predictor
ecoregions (Holdridge 1947). Plotting mean annual
biotemperature against mean annual precipitation variables. No parametric assumptions are necessary, but
and potential evapotranspiration ratio places any loca- the probability distribution (e.g., binomial, Poisson,
tion into one of a set of predetermined equal-area Gaussian) of the response variable must be speci ed.
Generalized Linear Modeling (GLM) is a special case of
hexagons, which are assigned ecoregion names a priori.
GAMs in which predictors are parameterized instead of
Homogeneity is not controlled, because several hexa-
being smoothed (McCullagh and Nelder 1997).
gons could represent a single ecoregion type. Lugo and
S41
Potential of Quantitative Methods for Ecoregions
Generalized additive modelings may be biased if freshwater shes. Iverson and Prasad (1998) used RTA
they are tted using presence-only datasets (e.g., mu- to generate predicted ranges for 80 tree species in the
seum collection localities). Because of their sequential eastern United States following climate change. Prince
nature, GAMs are poor at capturing interactions and Steininger (1999) used RTA with six forcing vari-
among predictor variables. A crucial step is the selec- ables, including rainfall, temperature, and photosyn-
tion of an appropriate level of spatial smoothing for a thetically active radiation, to stratify sampling in the
predictor. Relationships between predictor variables Large Scale Biosphere Atmosphere Experiment in
and the response variable are empirical rather than Amazonia.
mechanistic or process based. Therefore, GAMs tted Several quantitative environmental envelope-based
to data in a small region generally do not extrapolate methods have been developed to ascertain habitat
well across space or time. suitability. Busby (1991) developed BIOCLIM, which
Generalized additive modeling was used by Austin uses a simple bounding hyperbox method to capture
and others (1990) to model the niches of ve Eucalypt species occurrences in data space. Although widely
species. Overton and others (2001) used GAM to used, BIOCLIM overpredicts when species distribution
generate geographic frameworks in New Zealand for is in uenced by a combination of environmental pre-
the purposes of ecosystem management and sustain- dictors rather than by each one individually (Carpenter
ability. Overton and others (2000), Leathwick (2001), and others 1993). Walker and Cocks (1991) produced
and Lehmann and others (2002a) have used GAM HABITAT, which attempts to form a convex envelope
repeatedly to predict occurrence of many species, more tightly containing all species occurrences than
building up estimates of community structure and the simple rectilinear envelope of BIOCLIM. Carpen-
biodiversity. They followed a Predict First, Classify La- ter and others (1993) found that BIOCLIM overpre-
ter (PFCL) paradigm. Brooker and others (2002) dicted and HABITAT underpredicted habitat, and they
demonstrated the utility of ecoregions in epidemiology proposed a new method, DOMAIN, based on similarity
and human health by using logistic regression model- with all occupied points using Gower s metric. This
ing to quantitatively delineate schistosomiasis regions metric is the arithmetic mean of the differences be-
in Africa. Lehmann and others (2002b) created Gen- tween the two points in each dimension, after being
eralized Regression Analysis and Spatial Prediction standardized by the range to equalize the contribution
(GRASP), which formalizes the regression modeling from each predictor. Hirzel and Arlettaz (2003) point
approach to species distribution modeling using GAM. out that the Gower metric does not consider the den-
See extensive reviews of GAM in Guisan and others sity of occupied points and is, therefore, subject to
(2002) and Guisan and Zimmerman (2000). in uence by outliers.
Classi cation and Regression Trees (CART) and Hirzel and others (2002) describe Ecological Niche
Regression Tree Analysis (RTA) for categorical and Factor Analysis (ENFA), which does not require true
continuous response variables, respectively, have been absence data. ENFA discriminates environments occu-
used for both regionalization (Stoms and Hargrove pied by species from all environmental combinations
2000) and geographic range prediction (Iverson and occurring within a larger study area. ENFA models the
others 1999). A regression tree is a binary decision tree environmental niche relative to some set of environ-
in which branching at each step is de ned by test cri- mental variance found within a larger study area. The
teria involving a single best predictor variable, which range of conditions represented in the larger area
can be continuous or categorical. CART and RTA chosen therefore constrains the de nition of the
over t a tree on a training sample, which then has environmental niche envelope based on its context.
almost as many terminal nodes as there are training Hirzel and others (2001) tested ENFA and GLM
observations. Nodes are then pruned from the tree or against three virtual species that were spreading, at
shrunk (at the cost of decreased accuracy) to achieve equilibrium, and overabundant. GLM was badly af-
generality. CART and RTA choose the locally best dis- fected for the spreading species, but produced slightly
criminatory feature at each stage in the divisive process better results than ENFA when the species was over-
rather than the globally best discriminator (Stockwell abundant. Both methods produced equivalent results
and Noble 1992), and they enforce a sequential uni- when the virtual species was at equilibrium.
variate model rather than a true multivariate approach. Zaniewski and others (2002) compared GAMs with
White and others (1999) used RTA to predict bird presence-only data to GAMs using computer-generated
pseudoabsences and ENFA models. By using the
occurrence data in Oregon in terms of 10 environ-
mental variables, and Rathert and others (1999) found same presence data for all models, absence data were
environmental correlates with richness of Oregon isolated as the varying factor, allowing different tech-
S42 W. W. Hargrove and F. M. Hoffman
niques for modeling presence-only and pseudoabsence clustering is computationally intensive, so the assem-
data to be compared. Although presence-only GAMs blage to be classi ed must be limited to relatively few
predicted individual species distributions more accu- objects. Nonhierarchical clustering provides a single,
rately than ENFA, they were less effective than ENFA in user-speci ed level of division into groups; however, it
highlighting biodiversity hot spots from the sum- can be used to classify a large number of objects be-
ming of species predictions. Engler and others (2004) cause it does not divide exhaustively.
found that using ENFA-weighted pseudoabsences en- Cluster analysis can use occurrence data from a set
hanced the quality of GLM-based potential distribution of geographic sites to identify biotic communities that
maps for rare and endangered species. coexist in space. Overpeck and others (1985) clustered
Stockwell and Noble (1992) described a Genetic fossil pollens through time to identify modern analogs
Algorithm for Ruleset Prediction (GARP), which uses a for ancient vegetation assemblages retrieved from
genetic algorithm and four types of rule to inductively cores. Campbell and McAndrews (1991) used a cluster
develop a rule set for predicting a geographic range analysis to group 33 lakes in southern Ontario. They
from a set of presence-only data for a particular species. described community changes in response to the Little
Starting with an initial set of approximate rules, a ge- Ice Age within the region by grouping similar pollen
netic algorithm modi es them in ways that might (or diagrams.
might not) lead to an increase in predictive power.
Multivariate Clustering to Form Homogeneous
Randomize, mutate, and concatenate operators alter
Regions
individual rules in a rule set, and optimized sets of
Several investigators have recognized the potential
rules with the best predictive ability are retained.
of geographic multivariate clustering for delineating
GARP predicts a slightly different geographic range
homogeneous regions objectively within small maps
using each optimized rule set. These alternative ranges
(Belbin 1993; Omi and others 1979; Host and others
are usually overlaid geographically and used as a single
1996; Bunce and others 1996). Host and others (1996)
pseudosuitability map (Peterson and others 2003).
used clustering to establish separate climatic and
Because GARP can suffer problems with repeatability,
lack of absence data, and variable prior proportions, physiographic regions for northern Wisconsin, but
then combined them using a simple GIS overlay.
Stockwell and Peters (1999) developed a web-based
Environmental characteristics have been clustered to
system that eliminates these and other potential sour-
produce uniform regions of geology (Harff and Davis
ces of error. Peterson and Cohoon (1999) found that
1990), regions of uniform crop yield (Lark 1998), and
GARP predictions were sensitive to the number of
regions of constant soil fertility (Carter 1997). Bernert
environmental variables that are included. Including
and others (1997) used cluster analysis to further sub-
ve of eight environmental layers was necessary to
divide an existing expert-derived ecoregion, the Wes-
avoid broad variance in predicted ranges for three
species of birds. GARP can utilize museum collection tern Corn Belt Plains, and Krohn and others (1999)
used geographic clustering to create hierarchical bio-
data and has been used to estimate the ultimate po-
physical regions of Maine at 21 km resolution. Soriano
tential distributions for invasive species (Peterson and
and Paruelo (1992) used normalized multidimensional
Vieglais 2001). Martinez-Meyer and others (2004) used
ordination of remotely sensed data to form homoge-
GARP to predict geographic distributions of 23 extant
neous biozones for Patagonia in southern Argentina.
mammal species reciprocally between the Last Glacial
Hessburg and others (2000) used TWINSPAN, a
Maximum and the present, suggesting that ecological
divisive hierarchical method to create groups of sub-
niches are relatively constant over time. Such longitu-
dinal evolutionary conservatism suggests that niche watersheds of the Columbia River Basin. Repeated
divisive classi cations of overlapping regions allowed
modeling can be used successfully to anticipate climate
construction of pedigree trees showing similar analysis
change effects on biodiversity.
ancestries among subregions. Region separation was
Quantitative Clustering of Biotic Assemblages evaluated using discriminant analysis and cross-valida-
Statistical clustering is the ordination and classi - tion. Hessburg and others (2000) prestrati ed using an
cation of multiple nonidentical objects into subgroups expert-derived regionalization and binned continuous
based on their similarity. Hierarchical clustering pro- ordinal variables before starting their quantitative
vides a series of divisions, based on some measure of process. Not all of their resultant regions nested within
similarity, into all possible numbers of groups, from even the largest expert-derived domains with which
one single group containing all objects, to potentially they were initially constrained. Sequentially divisive
as many groups as there are objects. Hierarchical methods are unlikely to result in equal-variance re-
S43
Potential of Quantitative Methods for Ecoregions
gions, particularly if some regions are subsequently deviation. Because the plotted location of map cells in
recombined. data space pinpoints the combination of environmen-
Jensen and others (2001) used 19 indirect biophys- tal variables within that map cell, two map cells that are
ical variables to hierarchically classify subwatersheds in plotted close to one another in data space will have
the Columbia River Basin using agglomerative cluster- similar mixtures of environmental conditions and are
ing with Ward s method. Subsequent analysis of vari- likely to be classi ed into the same ecoregion cluster.
ance (ANOVA) showed that the hydrologic subregions Thus, similarity is coded as separation distance in
produced were highly signi cant in explaining Forest environmental data space.
Service watershed and stream management hazard We use the iterative k-means algorithm of Hartigan
ratings, although no comparison was made with the (1975), which begins with a user-specified number of
regionalization of Hessburg and others (2000) or with ecoregion clusters, k, into which the map cells are to be
expert-derived regions. grouped. All map cells are examined sequentially to
Zhou and others (2003) created objective ecore- find the most widely separated set of cells that will
provide k initial seed centroids, one for each of the
gions of Nebraska using an agglomerative hierarchical
clustering procedure based on multitemporal satellite desired k cluster groups. In a single iteration, each map
data, along with climate and soil information. Their cell is assigned to the closest (i.e., environmentally
aggregation procedure combined 2024 polygons to most similar) centroid. At the end of the iteration,
form 66 and 23 hierarchical regions. Because only once all map cells are assigned to a centroid, the
spatially adjacent regions are merged, island and bar- coordinates of all map cells within each group are
rier features (i.e., water bodies and rivers) require averaged to produce a new, adjusted centroid for each
human intervention during processing to prevent lin- cluster, and another iteration of assigning map cells to
gering arbitrary separations. these new centroids begins. The iterative process of
Leathwick and others (2003) created environ- classifying map cells and adjusting centroid locations
mental domains at 1 km resolution for New Zealand continues until fewer than a predetermined number of
using a 2-stage multivariate classi cation based on 10 map cells change cluster assignments during an itera-
climatic and landform variables affecting plant physi- tion. After the process has converged on a particular
ological processes. They rst used nonhierarchical grouping scheme, the k ecoregions have been statisti-
clustering to produce 350 geographic groups, then cally defined.
used sequential agglomerative clustering to obtain 20 Once cluster assignments have stabilized, map cells
nal domains. In both stages, they used the Gower are reassembled in geographic space, retaining their
metric as a similarity measure, which is sensitive to ecoregion classi cations. Although geographic coordi-
outliers (Hirzel and Arlettaz 2003). They presented a nates are not used directly in the classi cation, ecore-
tree showing the similarity of the 20 nal domains, gions tend to be geographically cohesive because of the
each of which were intuitively recognizable. spatial autocorrelation that is usually present in the
We developed a supercomputer-based multivariate environmental data. Because of the Euclidean assign-
statistical clustering algorithm to de ne ecoregions ment method, the k-means algorithm tends to fit
within extensive, high-resolution maps containing globular clusters of equal size in data space. Thus, all
many multivariate descriptors (Hargrove and others large ecoregions share a similar upper limit on within-
2001) (Figure 1). The user can specify the number of group variance and have a similar maximum radius
clustered ecoregions that result from the process, around each centroid. Although other clustering
making it possible to divide the map into a few large, methods (i.e., Ward s, average, single, or complete
coarsely-defined ecoregions or a larger number of linkage) are available, the uniform heterogeneity
small, highly specified ones. across ecoregions provided by k-means prevents the
The nonhierarchical algorithm consists of a revers- creation of side-by-side ecoregions that have vastly dif-
ible transformation between two realms: one in two- ferent within-region variance.
dimensional geographic map space and one in multi- We call this empirical process Multivariate Geo-
dimensional data space. Normalized variable values graphic Clustering (MGC) and have implemented it in
from each map raster cell are used as coordinates to a parallel algorithm coded in C using the Message
plot each map cell in an environmental space with as Passing Interface (MPI). Our code is dynamically load
many axes as there are multivariate environmental balancing and fault tolerant and performs both initial
dimensions. Normalization gives environmental seed- nding and iterative cluster assignment in parallel
parameters measured in different units equal spacing (Hoffman and Hargrove 1999). Individual nodes
by establishing a mean of zero and a unit standard independently classify subsets of cells, then combine
S44 W. W. Hargrove and F. M. Hoffman
Figure 1. The Multivariate Spatio-Temporal Clustering procedure involves a transformation (moving clockwise from upper
left) from geographic space (green) to data space (blue) and back. Normalized values of multivariate conditions in each map
cell are used as coordinates to locate each map cell in an abstract data space. Although only three axes are shown, the process
considers many multivariate characteristics. An iterative k-means clustering procedure assigns each cell to the closest of k
centroids. At the end of each iteration, centroid positions are recomputed. After convergence, cells are reassembled in geo-
graphic space, colored by their final cluster assignments. Each resultant quantitative ecoregion contains roughly equal envi-
ronmental heterogeneity.
results at the end of an iteration. As a result, the allowing a uniform classification structure to emerge
quantitative ecoregionalization process is not compu- from the data.
tationally limited. Strong correlations among input variables will affect
Ecoregions determined using MGC are self- clustering results. For example, elevation might be
describing, in that the coordinates of the nal cent- correlated with temperature and other climatic vari-
roids quantitatively de ne the synoptic conditions for ables. Each input map should be selected or designed
each ecoregion. Nominative multivariate conditions to contain unique information in order to preserve
within a particular ecoregion are described by the N orthogonality in data space. Even strongly correlated
coordinates of its centroid. The technique is paramet- inputs can be clustered by rst performing a Principal
ric only in that means are used to calculate each new Components Analysis (PCA) and then clustering in a
centroid. Rather than imposing a preconceived exter- data space formed by the PCA axes. Patterns resolved
nal grouping upon the map, the variance structure by MGC have proven robust to even strong correlations
present in the environmental conditions is used, among a few input variables.
S45
Potential of Quantitative Methods for Ecoregions
factors, and climatic factors. The edaphic factors were
(1) plant-available water capacity, (2) soil organic
matter, (3) total Kjeldahl soil nitrogen, and (4) depth
to a seasonally high water table. The climatic factors
were (1) mean precipitation during the growing sea-
son, (2) mean solar insolation during the growing
season, (3) degree-day heat sum during the growing
season, and (4) degree-day cold sum during the
nongrowing season. The growing season was defined
by the frost-free period between mean day of first and
last frost each year. A map for each of these charac-
teristics was generated from best available data at a 1-
km resolution for input into the clustering process.
Each of the input maps contained more than 7.8
million cells. Such map, data, and ecoregion resolu-
tion surpasses that usually accomplished by ecoregion
experts using qualitative methods. These maps appear
to capture the ecological relationships among the
nine input variables (Figure 2a). More recently, we
have added new variables and divided the United
States into as many as 5000 ecoregions. Twenty-five
environmental factors, including elevation, mean and
extremes of annual temperature, mean monthly pre-
cipitation, soil nitrogen, organic matter and water
capacity, frost-free days, soil bulk density and depth,
and solar aspect and insolation were included (Har-
Figure 2. (A) The 3000 most different quantitative ecore-
grove and Hoffman 2003).
gions in the United States based on nine variables, including
precipitation, solar input, elevation, depth to water table, soil
Visualizing Ecoregion Similarity Using
nitrogen and organic matter, soil water-holding capacity,
heating degree-days during the growing season, and cooling
Similarity Colors
degree-days during the nongrowing season, colored ran-
Randomly colored ecoregion maps emphasize the
domly. (B) When PCA is performed on the nine variables and
the first three scores are assigned to red, green and blue, the location of borders between ecoregions. However,
color of each ecoregion indicates the relative mix of the nine ecologists might also wish for some indication of how
environmental conditions inside each ecoregion. Red is different the mixture of environmental conditions is
physiographic position (i.e., low precipitation, high solar
across the border between neighboring ecoregions.
insolation, high elevation, and deep water table). Green is
Because the nal location of the cluster centroid is, by
plant nutrients (i.e., high soil N, organic matter, and
de nition, the most centrally located point inside each
available water). Blue is temperature (i.e., few degree-days
cluster, the data space coordinates of the centroid
heat and many degree-days cool). Shown in these Similarity
provide a description of the average ecological condi-
Colors, the borders between individual ecoregions disappear
tions in this cluster ecoregion. Likewise, differences
and the map now shows regional scale gradients in environ-
between the centroid coordinates from two ecoregions
mental conditions.
quantify the differences between the average environ-
ments found in each cluster ecoregion.
We devised a statistical coloring scheme to visualize
Multivariate Geographic Clustering Across
similarities among environments within different eco-
Space
regions. If we use a PCA, either before or after MGC, to
condense a larger number of raw environmental
Initially, we performed a number of empirical
regionalizations for the conterminous United States at variables into orthogonal principal component axes,
1 km resolution, dividing the United States into as we can map the rst, second, and third principal
many as 3000 distinct ecoregions (Figure 2A), (Har- component scores to a red green blue (RGB) color
grove and Luxmoore 1998). We included nine char- triplet. In this way, the combination of values for the
acteristics from three categories elevation, edaphic coordinates for each cluster centroid are used to
S46 W. W. Hargrove and F. M. Hoffman
specify a unique color for that ecoregion that indicates
the relative mixture of each environmental factor.
Ecoregions containing similar environmental combi-
nations are colored similarly.
The ecoregion map using Similarity Colors appears
strikingly different from that using random colors.
Individual ecoregion borders disappear and the Sim-
ilarity Colors reveal smooth gradients that re ect the
dominant suites of variables affecting environments in
each region of the country (Figure 2B). The red
Southwest is dominated by physiographic factors such
as low precipitation, high solar radiation, and a deep
water table. The blue Northeast is dominated by
temperature factors, and the green Southeast is
Figure 3. Borders between adjacent quantitative ecoregions
dominated by fertile soils. The upper Midwest is high
can be characterized as sharp or gradual using representa-
in all factors, but is light blue because of the cold
tiveness contours. Georgia is at the right, Alabama is at the
continental winter. The Pacific Northwest and the
left, and the Florida panhandle is at the bottom. Ecoregions
Central California valley are light green, indicating are shown in Similarity Colors, as in Figure 2B. The repre-
favorable conditions for plants. It is not possible to sentativeness elevation of each cell is calculated by measuring
code maps by Similarity Colors when ecoregions are the Euclidean distance in data space from the cell to the
created qualitatively. centroid of the quantitative ecoregion to which the cell was
assigned. Representativeness contours are closely spaced and
parallel adjacent to sharp borders, but meander near gradual
Characterizing Borders Between Ecoregions borders.
Borders between ecoregions can be sharp, forming
distinct ecotones. More commonly, however, gradual
two distinct sharpness properties. We used contour
transitions cause edges to be indistinct, and it is dif -
lines of equal representativeness to visualize the
cult to locate a line of demarcation between distinct
sharpness characteristics of borders between ecore-
ecoregions (Bailey 1983). We have termed this type of
gradual border an ecopause (Hargrove and Hoff- gions (Hargrove and Hoffman 1999). Closely spaced
representativeness contours re ect steep edges and,
man 1999). Indeed, a border can begin at one geo-
therefore, a sharp ecotone. On the other hand, widely
graphic location as an ecotone, and then transform
spaced representativeness contours indicate gradually
slowly along its length into an ecopause. Approaches to
sloping representativeness, characteristic of an indis-
distinguish these different types of border have in-
tinct ecopause. Representativeness contour lines have
cluded fuzzy set theory (Leung 1987; Lark 1998) and
the exibility to represent mixed gradual/sharp bor-
wavelet analysis (Csillag and others 2001; Csillag and
ders, as well as borders whose characteristics change
Kabos 2002), but no single method has been widely
along their length.
adopted.
Figure 3 shows the representativeness surface for
Fundamental to characterizing the sharpness of
southwest Georgia, Alabama, and northern Florida.
borders is quantifying how representative a particular
The topography of each cell is obtained from the
location is of its parent ecoregion. In MGC, the
Euclidean distance from that cell s location in envi-
Euclidean distance from each cell to its centroid
ronmental data space to the centroid of its ecoregion.
measures its deviation from the cluster norm. Cells
Cluster membership is shown in Figure 3 as the Simi-
close to their centroids are more representative of
larity Colors of each cell. The random orientation and
their cluster ecoregions than cells far from their
meandering character of the contours near the Coastal
centroids in environmental space (Belbin 1993). If we
depict these representativeness values as elevations, Plain and Piedmont ecoregions in southern Alabama
clearly indicate that this border is an ecopause (i.e.,
we can create a continuous surface whose height in-
environmental gradients are relatively unchanging).
versely corresponds to the representativeness of the
On the other hand, the closely spaced, parallel contour
cell at that geographic location (Hargrove and Hoff-
lines separating the Piedmont from the Ridge-and-
man 1999).
Valley in northern Alabama reveal this border as a
Because the edge properties are calculated for each
sharp ecotone.
pair of adjacent clusters, each side of every border has
S47
Potential of Quantitative Methods for Ecoregions
Figure 4. Maps of similarity to any selected quantitative ecoregion can be produced. The Euclidean distance in data space from
the centroid of each ecoregion to the centroid of the chosen region is calculated. Ecoregions with closer centroids are more
similar and are colored darker gray. The Everglades ecoregion was selected from this 1000-ecoregion map based on 25 primary
environmental variables, so that the map shows the quantitative degree of Everglades-ness across the map.
Abrupt changes in Similarity Colors are accompa- the degree of similarity of all ecoregions to the selected
nied by the numerous parallel contours of an ecotone, ecoregion can be mapped.
whereas subtle color changes are accompanied by the Thus, maps can be drawn that show the degree of
meandering contours of an ecopause. Close represen- innate similarity between a particular selected ecore-
tativeness contours create thick black lines where eco- gion and the rest of the map. For example, starting
region borders are sharp ecotones. Representativeness with the 1000 most different ecoregions based on the
contours combine the traditional notion of exclusive 25 primary environmental factors described earlier, we
produced a map of Everglades-ness that shows how
membership of each location in a single ecoregion
with a quantitative measure of goodness of t or similar other regions are to the Florida Everglades
belonging. (Figure 4). Variables considered include elevation,
temperature, precipitation, soil characteristics, and
solar inputs (Hargrove and Hoffman 2003). Darker
Quantifying the Similarity of Ecoregions
areas are most similar to the selected ecoregion. The
When ecoregions are quantitatively derived, one can Okefenokee Swamp in Georgia, the Great Dismal
select a single ecoregion of interest and then produce a Swamp in Virginia, the Mississippi Delta, and the Wis-
consin/Minnesota Land of a Thousand Lakes all
sorted list of the similarity of all other ecoregions to the
rank high in their degree of Everglades-ness. These
one selected. The chosen ecoregion establishes an
origin in data space and, using the Euclidean distance comparative representativeness maps quantify the
from this origin to the centroid of every other ecore- ecoregion comparisons that ecologists have always
gion, pairwise similarity measures can be calculated. wanted to make but, using traditional expertise-based
Coding these pairwise similarity values as gray levels, ecoregions, could only subjectively estimate.
S48 W. W. Hargrove and F. M. Hoffman
Quantitative Ecoregions as a Basis for Networks of installations like LTER and AmeriFlux
represent signi cant investments of research capital.
Sampling Network Design
Sites in most national-scale networks might not be lo-
cated by design, but, instead, by opportunity or logistic
If we invert the quantitative comparison concept to
convenience. Network analysis, as an outgrowth of the
consider nonrepresentativeness in the context of an
quantitative treatment of ecoregions, provides objec-
existing network of sites or sample locations, we can
tive guidance about the design, performance, and
quantify how well a particular established network is
modi cation of such networks and can improve on
representative of a larger map that contains it. A net-
more idiosyncratic design approaches.
work in this sense consists of a geographic constellation
of sites or facilities or can simply represent locations
where samples have been taken. Network analysis
shows how well the sampled environments represent Statistical Modeling of Environmental Niche
the rest of the map and identi es the best locations for
Envelopes
new sites or installations. The best location for an
A variation of cluster-based ecoregionalization uses
additional site will be in places that are the least well
Hutchinson s theoretical underpinnings to forecast a
represented by the network of existing sites.
species geographic range in new locations or under
Instead of a one-to-one centroid comparison like
Everglades-ness, network analysis entails a one-to- altered environmental conditions (Hargrove and
Hoffman 2000). First, each cell within the species
many centroid comparison in data space. To quantify
environmental range is located in a multidimensional
network coverage, we determine how different each
environmental space. Rather than specifying the
ecoregion is from the most similar network site or
number of ecoregion groups desired, variance-based
sample. For each ecoregion in the map, we nd the
clustering is used to de ne as many xed-radius, equal-
Euclidean distance in data space to the single closest
variance clusters as necessary to ensure that each of the
ecoregion that contains a site from the network. As
cells occupied by the species is contained within at least
earlier, this distance is coded to a gray level. Unlike
the Everglades-ness maps, however, darker areas one cluster. All clusters, when considered together as
multidimensional volumetric pixels, form a model of
represent areas that are poorly represented by the
the location and shape of the niche within data space.
existing network. Because this method quanti es
Selection of the radius (variance) for the clusters
coverage or presence of sites, sites will always sit
controls the resolution with which the niche envelope
within well-represented ecoregions, which will be col-
is de ned and serves to ll gaps within the Hutchin-
ored white.
sonian hypervolume.
Maps showing the geographic areas represented by
If the current geographic range data include some
each individual site can be generated, and importance
surrogate measures of tness under each combination
values can be calculated for each site based on the
of environmental conditions, this information can be
marginal representation it adds to the network.
attached to the clusters de ning the niche hypervo-
Quantifying the contribution of each site to network
lume. Hypervolume de nitions are used to make geo-
representation can minimize the impact of site elimi-
graphic range predictions for new areas by projecting
nation on representation. Finally, a network with a gi-
all map cells in the area into standardized environment
ven number of sites can be designed that is
space, then testing each cell to determine if it is within
theoretically optimal, having the highest possible rep-
one of the clusters that de ne the niche hypervolume.
resentation of environmental conditions on a map
The mean surrogate tness associated with the closest
(given that number of underlying ecoregions).
cluster centroid containing a cell is used to predict the
When submitted to a network analysis based on an
tness of the species within that cell location in the
ecoregionalization into the 2000 most different ecore-
new geographic range.
gions based on 25 environmental factors, the National
We de ned niche hypervolume models for loblolly
Science Foundation s Long-Term Ecological Research
pine, Pinus taeda L., and sugar maple, Acer saccharum
(LTER) study site network indicates that additional
Marsh., in terms of the 25 environmental condition
LTER sites in the Olympic peninsula and the Klamath,
gradients of climatic, physi