Post Job Free
Sign in

Data Environmental

Location:
Oak Ridge, TN
Posted:
November 15, 2012

Contact this candidate

Resume:

DOI: **.****/s*****-***-****-*

Potential of Multivariate Quantitative Methods for

Delineation and Visualization of Ecoregions

WILLIAM W. HARGROVE* from the limitations of human subjectivity, making possible a

FORREST M. HOFFMAN new array of ecologically useful derivative products. A

Environmental Sciences Division red green blue visualization based on principal components

Computer Science and Math Division analysis of ecoregion centroids indicates with color the rel-

Oak Ridge National Laboratory ative combination of environmental conditions found within

P. O. Box 2008, M.S. 6407 each ecoregion. Multiple geographic areas can be classified

Oak Ridge, Tennessee 37831-6407 into a single common set of quantitative ecoregions to pro-

vide a basis for comparison, or maps of a single area through

ABSTRACT / Multivariate clustering based on fine spatial time can be classified to portray climatic or environmental

resolution maps of elevation, temperature, precipitation, soil changes geographically in terms of current conditions.

characteristics, and solar inputs has been used at several Quantified representativeness can characterize borders be-

specified levels of division to produce a spectrum of tween ecoregions as gradual, sharp, or of changing char-

quantitative ecoregion maps for the conterminous United acter along their length. Similarity of any ecoregion to all other

States. The coarse ecoregion divisions accurately capture ecoregions can be quantified and displayed as a repre-

intuitively-understood regional environmental differences, sentativeness map. The representativeness of an existing

whereas the finer divisions highlight local condition spatial array of sample locations or study sites can be

gradients, ecotones, and clines. Such statistically generated mapped relative to a set of quantitative ecoregions, sug-

ecoregions can be produced based on user-selected gesting locations for additional samples or sites. In addition,

continuous variables, allowing customized regions to be the shape of Hutchinsonian niches in environment space can

delineated for any specific problem. By creating an objective be defined if a multivariate range map of species occurrence

ecoregion classification, the ecoregion concept is removed is available.

Ecoregions are designed to help users visualize and using human expertise in a qualitative, weight-of-evi-

understand similarities across complex multivariate dence approach (McMahon and others 2001).

environmental factors by grouping areas into like cat- It is not our intention to add to this debate. These

egories. The basis for such groupings has many con- conceptual issues are well represented in this special

tentious conceptual underpinnings. Debate abounds as issue and elsewhere. Rather, we sidestep the question

to whether ecoregions should be specialized for a as to whether ecoregions are computable and demon-

particular use or general purpose, spatially contiguous strate the rami cations of quantitatively deriving eco-

versus disjunct, nestable versus nonhierarchical, and regions. We argue that the spate of ancillary products

whether ecoregions can be defensible as units of resulting from the quantitative treatment of ecoregions

management, legislation, or even ecological triage enhances and expands the utility of the ecoregion

(Omernik 1995, 2003; Overton and others 2002; concept and makes substantial new contributions to

Leathwick and others 2003). Supreme among these niche modeling, network and sample design, change

issues, however, is the question of whether ecoregions detection, and conservation.

can (or should) be delineated using quantitative sta- Regionalizations are models, whether quantitatively

tistical methods or whether they can only be drawn or qualitatively derived. Quantitative models are, how-

ever, more explicit, repeatable, transferable, and

defensible than subjective models based on human

KEY WORDS: Clustering; Climate change; Ecotone; Environmental expertise. This transferability and repeatability makes

envelope; Fences; Gradient; Network; Niche; Pre- quantitative models more objective than their qualita-

serve design; Range; Representativeness; Similarity;

tive counterparts. Human experts might be able to

Time series

rationally defend drawing a particular borderline be-

tween ecoregions, but they might be unable to eluci-

Published online February 25, 2005.

date the method used to place it at that precise

*Author to whom correspondence should be addressed; email:

location. Quantitative ecoregionalization techniques

***@****.***.****.***

2005 Springer Science+Business Media, Inc.

Environmental Management Vol. 34, Suppl. 1, pp. S39 S60

S40 W. W. Hargrove and F. M. Hoffman

are not perfectly objective, because they still require others (1999) recently revisited the Life Zone approach

subjective ecological expertise in the choice of data for global vegetation, finding it adequate for forests, but

layers to include and in the interpretation of the of limited utility for grasslands, shrublands, and non-

resulting ecoregions. Nevertheless, the pursuit of a vegetated lands.

fully describable quantitative method for ecoregional- Kohonen s (1982, 1995) Self-Organizing Maps

ization is a desirable goal, whether as a replacement or (SOMs) use a two-layer neural network to divide a

an augmentation to qualitative, expert-opinion-based multivariate map into ecoregions. During training,

approaches. neurons are reinforced within a gradually shrinking

The ability to explicitly select and control the input network neighborhood around the best predictors

variables allows quantitative regionalizations to be cus- (Hung 1993). Using the rst ve principal components

tomized for speci c uses. Using quantitative methods, of seasonal averages of climate data from 18 stations,

special-purpose ecoregions can be created speci cally Malmgren and Winter (1999) used a one-dimensional

for particular uses. Such customized regionalizations SOM to divide Puerto Rico into four natural climatic

can provide additional discrimination in places where zones. It might be dif cult, however, to transfer the

learning from one neural network into a form that

popular generalized ecoregion schemes might miss

particular features of importance. can be used by others.

In qualitative ecoregionalization, human experts

Quantitative Ecoregions for Single Species

might intentionally adjust the weighting of particular

Mapping the geographic range of a single species of

input layers in their mental model in particular spots

animal or plant is a special case of ecoregion delinea-

across the map, whereas quantitative methods usually

tion. Ecological niche theory is central to understanding

produce regionalizations in which all input variables

how environmental change affects species abundance

receive equal weighting. Such evenly weighted ecore-

patterns (Jackson and Overpeck 2000). Hutchinson

gion products provide, at the very least, an initial basis

(1957) conceived of the niche as a multidimensional

from which factor weights can be subsequently spatially

hypervolume with dimensions de ned by the envi-

modi ed by human expertise. Explicit spatial maps of

altered weights, if known, could be considered directly ronmental factors that in uence the tness of individ-

uals of that species.

as inputs into a fully quantitative model.

Although Hutchinson s conception was static, envi-

ronmental change involves temporal alteration of

Sampling the Toolbox of Quantitative Methods

combinations of niche variables. Hutchinson s niche

envelope assumes that a steady-state equilibrium with

A full review of quantitative methods that have been

present environmental conditions has allowed ade-

used in regionalization is beyond the scope of this

article. Although not exhaustive, brief treatment of quate time for perfect adaptation and exhaustive

major types of quantitative approach might provide migration to all parts of the potential range. Acclima-

tion to changing conditions and historical factors lim-

context, and counterbalance more typical subjective

iting geographic dispersal are not considered by most

approaches described elsewhere in this special issue.

quantitative range-prediction techniques [but migra-

Most quantitative approaches rely on Gleasonian rela-

tion can be simulated in a subsequent step (i.e., Pet-

tionships between environmental patterns and species

erson and others 2003)].

occurrence, but some have been tested directly using

Generalized Additive Modeling (GAM) uses regres-

species distribution data, which include Clementsian

effects of biotic interactions on geographic distribu- sion modeling to establish empirical relationships be-

tions. Strengths and weaknesses of each quantitative tween a response variable (i.e., presence of a particular

species at a location) and an individually smoothed set

method are highlighted.

of spatial predictor variables (Hastie and Tibshirani

The Holdridge Life Zone model, usually shown as a

1990). GAM additively calculates the component re-

set of hexagons arranged inside a triangular plot, was

sponse and can handle nonlinear and nonmonotonic

an early quantitative multivariate approach for de ning

relationships between the response and predictor

ecoregions (Holdridge 1947). Plotting mean annual

biotemperature against mean annual precipitation variables. No parametric assumptions are necessary, but

and potential evapotranspiration ratio places any loca- the probability distribution (e.g., binomial, Poisson,

tion into one of a set of predetermined equal-area Gaussian) of the response variable must be speci ed.

Generalized Linear Modeling (GLM) is a special case of

hexagons, which are assigned ecoregion names a priori.

GAMs in which predictors are parameterized instead of

Homogeneity is not controlled, because several hexa-

being smoothed (McCullagh and Nelder 1997).

gons could represent a single ecoregion type. Lugo and

S41

Potential of Quantitative Methods for Ecoregions

Generalized additive modelings may be biased if freshwater shes. Iverson and Prasad (1998) used RTA

they are tted using presence-only datasets (e.g., mu- to generate predicted ranges for 80 tree species in the

seum collection localities). Because of their sequential eastern United States following climate change. Prince

nature, GAMs are poor at capturing interactions and Steininger (1999) used RTA with six forcing vari-

among predictor variables. A crucial step is the selec- ables, including rainfall, temperature, and photosyn-

tion of an appropriate level of spatial smoothing for a thetically active radiation, to stratify sampling in the

predictor. Relationships between predictor variables Large Scale Biosphere Atmosphere Experiment in

and the response variable are empirical rather than Amazonia.

mechanistic or process based. Therefore, GAMs tted Several quantitative environmental envelope-based

to data in a small region generally do not extrapolate methods have been developed to ascertain habitat

well across space or time. suitability. Busby (1991) developed BIOCLIM, which

Generalized additive modeling was used by Austin uses a simple bounding hyperbox method to capture

and others (1990) to model the niches of ve Eucalypt species occurrences in data space. Although widely

species. Overton and others (2001) used GAM to used, BIOCLIM overpredicts when species distribution

generate geographic frameworks in New Zealand for is in uenced by a combination of environmental pre-

the purposes of ecosystem management and sustain- dictors rather than by each one individually (Carpenter

ability. Overton and others (2000), Leathwick (2001), and others 1993). Walker and Cocks (1991) produced

and Lehmann and others (2002a) have used GAM HABITAT, which attempts to form a convex envelope

repeatedly to predict occurrence of many species, more tightly containing all species occurrences than

building up estimates of community structure and the simple rectilinear envelope of BIOCLIM. Carpen-

biodiversity. They followed a Predict First, Classify La- ter and others (1993) found that BIOCLIM overpre-

ter (PFCL) paradigm. Brooker and others (2002) dicted and HABITAT underpredicted habitat, and they

demonstrated the utility of ecoregions in epidemiology proposed a new method, DOMAIN, based on similarity

and human health by using logistic regression model- with all occupied points using Gower s metric. This

ing to quantitatively delineate schistosomiasis regions metric is the arithmetic mean of the differences be-

in Africa. Lehmann and others (2002b) created Gen- tween the two points in each dimension, after being

eralized Regression Analysis and Spatial Prediction standardized by the range to equalize the contribution

(GRASP), which formalizes the regression modeling from each predictor. Hirzel and Arlettaz (2003) point

approach to species distribution modeling using GAM. out that the Gower metric does not consider the den-

See extensive reviews of GAM in Guisan and others sity of occupied points and is, therefore, subject to

(2002) and Guisan and Zimmerman (2000). in uence by outliers.

Classi cation and Regression Trees (CART) and Hirzel and others (2002) describe Ecological Niche

Regression Tree Analysis (RTA) for categorical and Factor Analysis (ENFA), which does not require true

continuous response variables, respectively, have been absence data. ENFA discriminates environments occu-

used for both regionalization (Stoms and Hargrove pied by species from all environmental combinations

2000) and geographic range prediction (Iverson and occurring within a larger study area. ENFA models the

others 1999). A regression tree is a binary decision tree environmental niche relative to some set of environ-

in which branching at each step is de ned by test cri- mental variance found within a larger study area. The

teria involving a single best predictor variable, which range of conditions represented in the larger area

can be continuous or categorical. CART and RTA chosen therefore constrains the de nition of the

over t a tree on a training sample, which then has environmental niche envelope based on its context.

almost as many terminal nodes as there are training Hirzel and others (2001) tested ENFA and GLM

observations. Nodes are then pruned from the tree or against three virtual species that were spreading, at

shrunk (at the cost of decreased accuracy) to achieve equilibrium, and overabundant. GLM was badly af-

generality. CART and RTA choose the locally best dis- fected for the spreading species, but produced slightly

criminatory feature at each stage in the divisive process better results than ENFA when the species was over-

rather than the globally best discriminator (Stockwell abundant. Both methods produced equivalent results

and Noble 1992), and they enforce a sequential uni- when the virtual species was at equilibrium.

variate model rather than a true multivariate approach. Zaniewski and others (2002) compared GAMs with

White and others (1999) used RTA to predict bird presence-only data to GAMs using computer-generated

pseudoabsences and ENFA models. By using the

occurrence data in Oregon in terms of 10 environ-

mental variables, and Rathert and others (1999) found same presence data for all models, absence data were

environmental correlates with richness of Oregon isolated as the varying factor, allowing different tech-

S42 W. W. Hargrove and F. M. Hoffman

niques for modeling presence-only and pseudoabsence clustering is computationally intensive, so the assem-

data to be compared. Although presence-only GAMs blage to be classi ed must be limited to relatively few

predicted individual species distributions more accu- objects. Nonhierarchical clustering provides a single,

rately than ENFA, they were less effective than ENFA in user-speci ed level of division into groups; however, it

highlighting biodiversity hot spots from the sum- can be used to classify a large number of objects be-

ming of species predictions. Engler and others (2004) cause it does not divide exhaustively.

found that using ENFA-weighted pseudoabsences en- Cluster analysis can use occurrence data from a set

hanced the quality of GLM-based potential distribution of geographic sites to identify biotic communities that

maps for rare and endangered species. coexist in space. Overpeck and others (1985) clustered

Stockwell and Noble (1992) described a Genetic fossil pollens through time to identify modern analogs

Algorithm for Ruleset Prediction (GARP), which uses a for ancient vegetation assemblages retrieved from

genetic algorithm and four types of rule to inductively cores. Campbell and McAndrews (1991) used a cluster

develop a rule set for predicting a geographic range analysis to group 33 lakes in southern Ontario. They

from a set of presence-only data for a particular species. described community changes in response to the Little

Starting with an initial set of approximate rules, a ge- Ice Age within the region by grouping similar pollen

netic algorithm modi es them in ways that might (or diagrams.

might not) lead to an increase in predictive power.

Multivariate Clustering to Form Homogeneous

Randomize, mutate, and concatenate operators alter

Regions

individual rules in a rule set, and optimized sets of

Several investigators have recognized the potential

rules with the best predictive ability are retained.

of geographic multivariate clustering for delineating

GARP predicts a slightly different geographic range

homogeneous regions objectively within small maps

using each optimized rule set. These alternative ranges

(Belbin 1993; Omi and others 1979; Host and others

are usually overlaid geographically and used as a single

1996; Bunce and others 1996). Host and others (1996)

pseudosuitability map (Peterson and others 2003).

used clustering to establish separate climatic and

Because GARP can suffer problems with repeatability,

lack of absence data, and variable prior proportions, physiographic regions for northern Wisconsin, but

then combined them using a simple GIS overlay.

Stockwell and Peters (1999) developed a web-based

Environmental characteristics have been clustered to

system that eliminates these and other potential sour-

produce uniform regions of geology (Harff and Davis

ces of error. Peterson and Cohoon (1999) found that

1990), regions of uniform crop yield (Lark 1998), and

GARP predictions were sensitive to the number of

regions of constant soil fertility (Carter 1997). Bernert

environmental variables that are included. Including

and others (1997) used cluster analysis to further sub-

ve of eight environmental layers was necessary to

divide an existing expert-derived ecoregion, the Wes-

avoid broad variance in predicted ranges for three

species of birds. GARP can utilize museum collection tern Corn Belt Plains, and Krohn and others (1999)

used geographic clustering to create hierarchical bio-

data and has been used to estimate the ultimate po-

physical regions of Maine at 21 km resolution. Soriano

tential distributions for invasive species (Peterson and

and Paruelo (1992) used normalized multidimensional

Vieglais 2001). Martinez-Meyer and others (2004) used

ordination of remotely sensed data to form homoge-

GARP to predict geographic distributions of 23 extant

neous biozones for Patagonia in southern Argentina.

mammal species reciprocally between the Last Glacial

Hessburg and others (2000) used TWINSPAN, a

Maximum and the present, suggesting that ecological

divisive hierarchical method to create groups of sub-

niches are relatively constant over time. Such longitu-

dinal evolutionary conservatism suggests that niche watersheds of the Columbia River Basin. Repeated

divisive classi cations of overlapping regions allowed

modeling can be used successfully to anticipate climate

construction of pedigree trees showing similar analysis

change effects on biodiversity.

ancestries among subregions. Region separation was

Quantitative Clustering of Biotic Assemblages evaluated using discriminant analysis and cross-valida-

Statistical clustering is the ordination and classi - tion. Hessburg and others (2000) prestrati ed using an

cation of multiple nonidentical objects into subgroups expert-derived regionalization and binned continuous

based on their similarity. Hierarchical clustering pro- ordinal variables before starting their quantitative

vides a series of divisions, based on some measure of process. Not all of their resultant regions nested within

similarity, into all possible numbers of groups, from even the largest expert-derived domains with which

one single group containing all objects, to potentially they were initially constrained. Sequentially divisive

as many groups as there are objects. Hierarchical methods are unlikely to result in equal-variance re-

S43

Potential of Quantitative Methods for Ecoregions

gions, particularly if some regions are subsequently deviation. Because the plotted location of map cells in

recombined. data space pinpoints the combination of environmen-

Jensen and others (2001) used 19 indirect biophys- tal variables within that map cell, two map cells that are

ical variables to hierarchically classify subwatersheds in plotted close to one another in data space will have

the Columbia River Basin using agglomerative cluster- similar mixtures of environmental conditions and are

ing with Ward s method. Subsequent analysis of vari- likely to be classi ed into the same ecoregion cluster.

ance (ANOVA) showed that the hydrologic subregions Thus, similarity is coded as separation distance in

produced were highly signi cant in explaining Forest environmental data space.

Service watershed and stream management hazard We use the iterative k-means algorithm of Hartigan

ratings, although no comparison was made with the (1975), which begins with a user-specified number of

regionalization of Hessburg and others (2000) or with ecoregion clusters, k, into which the map cells are to be

expert-derived regions. grouped. All map cells are examined sequentially to

Zhou and others (2003) created objective ecore- find the most widely separated set of cells that will

provide k initial seed centroids, one for each of the

gions of Nebraska using an agglomerative hierarchical

clustering procedure based on multitemporal satellite desired k cluster groups. In a single iteration, each map

data, along with climate and soil information. Their cell is assigned to the closest (i.e., environmentally

aggregation procedure combined 2024 polygons to most similar) centroid. At the end of the iteration,

form 66 and 23 hierarchical regions. Because only once all map cells are assigned to a centroid, the

spatially adjacent regions are merged, island and bar- coordinates of all map cells within each group are

rier features (i.e., water bodies and rivers) require averaged to produce a new, adjusted centroid for each

human intervention during processing to prevent lin- cluster, and another iteration of assigning map cells to

gering arbitrary separations. these new centroids begins. The iterative process of

Leathwick and others (2003) created environ- classifying map cells and adjusting centroid locations

mental domains at 1 km resolution for New Zealand continues until fewer than a predetermined number of

using a 2-stage multivariate classi cation based on 10 map cells change cluster assignments during an itera-

climatic and landform variables affecting plant physi- tion. After the process has converged on a particular

ological processes. They rst used nonhierarchical grouping scheme, the k ecoregions have been statisti-

clustering to produce 350 geographic groups, then cally defined.

used sequential agglomerative clustering to obtain 20 Once cluster assignments have stabilized, map cells

nal domains. In both stages, they used the Gower are reassembled in geographic space, retaining their

metric as a similarity measure, which is sensitive to ecoregion classi cations. Although geographic coordi-

outliers (Hirzel and Arlettaz 2003). They presented a nates are not used directly in the classi cation, ecore-

tree showing the similarity of the 20 nal domains, gions tend to be geographically cohesive because of the

each of which were intuitively recognizable. spatial autocorrelation that is usually present in the

We developed a supercomputer-based multivariate environmental data. Because of the Euclidean assign-

statistical clustering algorithm to de ne ecoregions ment method, the k-means algorithm tends to fit

within extensive, high-resolution maps containing globular clusters of equal size in data space. Thus, all

many multivariate descriptors (Hargrove and others large ecoregions share a similar upper limit on within-

2001) (Figure 1). The user can specify the number of group variance and have a similar maximum radius

clustered ecoregions that result from the process, around each centroid. Although other clustering

making it possible to divide the map into a few large, methods (i.e., Ward s, average, single, or complete

coarsely-defined ecoregions or a larger number of linkage) are available, the uniform heterogeneity

small, highly specified ones. across ecoregions provided by k-means prevents the

The nonhierarchical algorithm consists of a revers- creation of side-by-side ecoregions that have vastly dif-

ible transformation between two realms: one in two- ferent within-region variance.

dimensional geographic map space and one in multi- We call this empirical process Multivariate Geo-

dimensional data space. Normalized variable values graphic Clustering (MGC) and have implemented it in

from each map raster cell are used as coordinates to a parallel algorithm coded in C using the Message

plot each map cell in an environmental space with as Passing Interface (MPI). Our code is dynamically load

many axes as there are multivariate environmental balancing and fault tolerant and performs both initial

dimensions. Normalization gives environmental seed- nding and iterative cluster assignment in parallel

parameters measured in different units equal spacing (Hoffman and Hargrove 1999). Individual nodes

by establishing a mean of zero and a unit standard independently classify subsets of cells, then combine

S44 W. W. Hargrove and F. M. Hoffman

Figure 1. The Multivariate Spatio-Temporal Clustering procedure involves a transformation (moving clockwise from upper

left) from geographic space (green) to data space (blue) and back. Normalized values of multivariate conditions in each map

cell are used as coordinates to locate each map cell in an abstract data space. Although only three axes are shown, the process

considers many multivariate characteristics. An iterative k-means clustering procedure assigns each cell to the closest of k

centroids. At the end of each iteration, centroid positions are recomputed. After convergence, cells are reassembled in geo-

graphic space, colored by their final cluster assignments. Each resultant quantitative ecoregion contains roughly equal envi-

ronmental heterogeneity.

results at the end of an iteration. As a result, the allowing a uniform classification structure to emerge

quantitative ecoregionalization process is not compu- from the data.

tationally limited. Strong correlations among input variables will affect

Ecoregions determined using MGC are self- clustering results. For example, elevation might be

describing, in that the coordinates of the nal cent- correlated with temperature and other climatic vari-

roids quantitatively de ne the synoptic conditions for ables. Each input map should be selected or designed

each ecoregion. Nominative multivariate conditions to contain unique information in order to preserve

within a particular ecoregion are described by the N orthogonality in data space. Even strongly correlated

coordinates of its centroid. The technique is paramet- inputs can be clustered by rst performing a Principal

ric only in that means are used to calculate each new Components Analysis (PCA) and then clustering in a

centroid. Rather than imposing a preconceived exter- data space formed by the PCA axes. Patterns resolved

nal grouping upon the map, the variance structure by MGC have proven robust to even strong correlations

present in the environmental conditions is used, among a few input variables.

S45

Potential of Quantitative Methods for Ecoregions

factors, and climatic factors. The edaphic factors were

(1) plant-available water capacity, (2) soil organic

matter, (3) total Kjeldahl soil nitrogen, and (4) depth

to a seasonally high water table. The climatic factors

were (1) mean precipitation during the growing sea-

son, (2) mean solar insolation during the growing

season, (3) degree-day heat sum during the growing

season, and (4) degree-day cold sum during the

nongrowing season. The growing season was defined

by the frost-free period between mean day of first and

last frost each year. A map for each of these charac-

teristics was generated from best available data at a 1-

km resolution for input into the clustering process.

Each of the input maps contained more than 7.8

million cells. Such map, data, and ecoregion resolu-

tion surpasses that usually accomplished by ecoregion

experts using qualitative methods. These maps appear

to capture the ecological relationships among the

nine input variables (Figure 2a). More recently, we

have added new variables and divided the United

States into as many as 5000 ecoregions. Twenty-five

environmental factors, including elevation, mean and

extremes of annual temperature, mean monthly pre-

cipitation, soil nitrogen, organic matter and water

capacity, frost-free days, soil bulk density and depth,

and solar aspect and insolation were included (Har-

Figure 2. (A) The 3000 most different quantitative ecore-

grove and Hoffman 2003).

gions in the United States based on nine variables, including

precipitation, solar input, elevation, depth to water table, soil

Visualizing Ecoregion Similarity Using

nitrogen and organic matter, soil water-holding capacity,

heating degree-days during the growing season, and cooling

Similarity Colors

degree-days during the nongrowing season, colored ran-

Randomly colored ecoregion maps emphasize the

domly. (B) When PCA is performed on the nine variables and

the first three scores are assigned to red, green and blue, the location of borders between ecoregions. However,

color of each ecoregion indicates the relative mix of the nine ecologists might also wish for some indication of how

environmental conditions inside each ecoregion. Red is different the mixture of environmental conditions is

physiographic position (i.e., low precipitation, high solar

across the border between neighboring ecoregions.

insolation, high elevation, and deep water table). Green is

Because the nal location of the cluster centroid is, by

plant nutrients (i.e., high soil N, organic matter, and

de nition, the most centrally located point inside each

available water). Blue is temperature (i.e., few degree-days

cluster, the data space coordinates of the centroid

heat and many degree-days cool). Shown in these Similarity

provide a description of the average ecological condi-

Colors, the borders between individual ecoregions disappear

tions in this cluster ecoregion. Likewise, differences

and the map now shows regional scale gradients in environ-

between the centroid coordinates from two ecoregions

mental conditions.

quantify the differences between the average environ-

ments found in each cluster ecoregion.

We devised a statistical coloring scheme to visualize

Multivariate Geographic Clustering Across

similarities among environments within different eco-

Space

regions. If we use a PCA, either before or after MGC, to

condense a larger number of raw environmental

Initially, we performed a number of empirical

regionalizations for the conterminous United States at variables into orthogonal principal component axes,

1 km resolution, dividing the United States into as we can map the rst, second, and third principal

many as 3000 distinct ecoregions (Figure 2A), (Har- component scores to a red green blue (RGB) color

grove and Luxmoore 1998). We included nine char- triplet. In this way, the combination of values for the

acteristics from three categories elevation, edaphic coordinates for each cluster centroid are used to

S46 W. W. Hargrove and F. M. Hoffman

specify a unique color for that ecoregion that indicates

the relative mixture of each environmental factor.

Ecoregions containing similar environmental combi-

nations are colored similarly.

The ecoregion map using Similarity Colors appears

strikingly different from that using random colors.

Individual ecoregion borders disappear and the Sim-

ilarity Colors reveal smooth gradients that re ect the

dominant suites of variables affecting environments in

each region of the country (Figure 2B). The red

Southwest is dominated by physiographic factors such

as low precipitation, high solar radiation, and a deep

water table. The blue Northeast is dominated by

temperature factors, and the green Southeast is

Figure 3. Borders between adjacent quantitative ecoregions

dominated by fertile soils. The upper Midwest is high

can be characterized as sharp or gradual using representa-

in all factors, but is light blue because of the cold

tiveness contours. Georgia is at the right, Alabama is at the

continental winter. The Pacific Northwest and the

left, and the Florida panhandle is at the bottom. Ecoregions

Central California valley are light green, indicating are shown in Similarity Colors, as in Figure 2B. The repre-

favorable conditions for plants. It is not possible to sentativeness elevation of each cell is calculated by measuring

code maps by Similarity Colors when ecoregions are the Euclidean distance in data space from the cell to the

created qualitatively. centroid of the quantitative ecoregion to which the cell was

assigned. Representativeness contours are closely spaced and

parallel adjacent to sharp borders, but meander near gradual

Characterizing Borders Between Ecoregions borders.

Borders between ecoregions can be sharp, forming

distinct ecotones. More commonly, however, gradual

two distinct sharpness properties. We used contour

transitions cause edges to be indistinct, and it is dif -

lines of equal representativeness to visualize the

cult to locate a line of demarcation between distinct

sharpness characteristics of borders between ecore-

ecoregions (Bailey 1983). We have termed this type of

gradual border an ecopause (Hargrove and Hoff- gions (Hargrove and Hoffman 1999). Closely spaced

representativeness contours re ect steep edges and,

man 1999). Indeed, a border can begin at one geo-

therefore, a sharp ecotone. On the other hand, widely

graphic location as an ecotone, and then transform

spaced representativeness contours indicate gradually

slowly along its length into an ecopause. Approaches to

sloping representativeness, characteristic of an indis-

distinguish these different types of border have in-

tinct ecopause. Representativeness contour lines have

cluded fuzzy set theory (Leung 1987; Lark 1998) and

the exibility to represent mixed gradual/sharp bor-

wavelet analysis (Csillag and others 2001; Csillag and

ders, as well as borders whose characteristics change

Kabos 2002), but no single method has been widely

along their length.

adopted.

Figure 3 shows the representativeness surface for

Fundamental to characterizing the sharpness of

southwest Georgia, Alabama, and northern Florida.

borders is quantifying how representative a particular

The topography of each cell is obtained from the

location is of its parent ecoregion. In MGC, the

Euclidean distance from that cell s location in envi-

Euclidean distance from each cell to its centroid

ronmental data space to the centroid of its ecoregion.

measures its deviation from the cluster norm. Cells

Cluster membership is shown in Figure 3 as the Simi-

close to their centroids are more representative of

larity Colors of each cell. The random orientation and

their cluster ecoregions than cells far from their

meandering character of the contours near the Coastal

centroids in environmental space (Belbin 1993). If we

depict these representativeness values as elevations, Plain and Piedmont ecoregions in southern Alabama

clearly indicate that this border is an ecopause (i.e.,

we can create a continuous surface whose height in-

environmental gradients are relatively unchanging).

versely corresponds to the representativeness of the

On the other hand, the closely spaced, parallel contour

cell at that geographic location (Hargrove and Hoff-

lines separating the Piedmont from the Ridge-and-

man 1999).

Valley in northern Alabama reveal this border as a

Because the edge properties are calculated for each

sharp ecotone.

pair of adjacent clusters, each side of every border has

S47

Potential of Quantitative Methods for Ecoregions

Figure 4. Maps of similarity to any selected quantitative ecoregion can be produced. The Euclidean distance in data space from

the centroid of each ecoregion to the centroid of the chosen region is calculated. Ecoregions with closer centroids are more

similar and are colored darker gray. The Everglades ecoregion was selected from this 1000-ecoregion map based on 25 primary

environmental variables, so that the map shows the quantitative degree of Everglades-ness across the map.

Abrupt changes in Similarity Colors are accompa- the degree of similarity of all ecoregions to the selected

nied by the numerous parallel contours of an ecotone, ecoregion can be mapped.

whereas subtle color changes are accompanied by the Thus, maps can be drawn that show the degree of

meandering contours of an ecopause. Close represen- innate similarity between a particular selected ecore-

tativeness contours create thick black lines where eco- gion and the rest of the map. For example, starting

region borders are sharp ecotones. Representativeness with the 1000 most different ecoregions based on the

contours combine the traditional notion of exclusive 25 primary environmental factors described earlier, we

produced a map of Everglades-ness that shows how

membership of each location in a single ecoregion

with a quantitative measure of goodness of t or similar other regions are to the Florida Everglades

belonging. (Figure 4). Variables considered include elevation,

temperature, precipitation, soil characteristics, and

solar inputs (Hargrove and Hoffman 2003). Darker

Quantifying the Similarity of Ecoregions

areas are most similar to the selected ecoregion. The

When ecoregions are quantitatively derived, one can Okefenokee Swamp in Georgia, the Great Dismal

select a single ecoregion of interest and then produce a Swamp in Virginia, the Mississippi Delta, and the Wis-

consin/Minnesota Land of a Thousand Lakes all

sorted list of the similarity of all other ecoregions to the

rank high in their degree of Everglades-ness. These

one selected. The chosen ecoregion establishes an

origin in data space and, using the Euclidean distance comparative representativeness maps quantify the

from this origin to the centroid of every other ecore- ecoregion comparisons that ecologists have always

gion, pairwise similarity measures can be calculated. wanted to make but, using traditional expertise-based

Coding these pairwise similarity values as gray levels, ecoregions, could only subjectively estimate.

S48 W. W. Hargrove and F. M. Hoffman

Quantitative Ecoregions as a Basis for Networks of installations like LTER and AmeriFlux

represent signi cant investments of research capital.

Sampling Network Design

Sites in most national-scale networks might not be lo-

cated by design, but, instead, by opportunity or logistic

If we invert the quantitative comparison concept to

convenience. Network analysis, as an outgrowth of the

consider nonrepresentativeness in the context of an

quantitative treatment of ecoregions, provides objec-

existing network of sites or sample locations, we can

tive guidance about the design, performance, and

quantify how well a particular established network is

modi cation of such networks and can improve on

representative of a larger map that contains it. A net-

more idiosyncratic design approaches.

work in this sense consists of a geographic constellation

of sites or facilities or can simply represent locations

where samples have been taken. Network analysis

shows how well the sampled environments represent Statistical Modeling of Environmental Niche

the rest of the map and identi es the best locations for

Envelopes

new sites or installations. The best location for an

A variation of cluster-based ecoregionalization uses

additional site will be in places that are the least well

Hutchinson s theoretical underpinnings to forecast a

represented by the network of existing sites.

species geographic range in new locations or under

Instead of a one-to-one centroid comparison like

Everglades-ness, network analysis entails a one-to- altered environmental conditions (Hargrove and

Hoffman 2000). First, each cell within the species

many centroid comparison in data space. To quantify

environmental range is located in a multidimensional

network coverage, we determine how different each

environmental space. Rather than specifying the

ecoregion is from the most similar network site or

number of ecoregion groups desired, variance-based

sample. For each ecoregion in the map, we nd the

clustering is used to de ne as many xed-radius, equal-

Euclidean distance in data space to the single closest

variance clusters as necessary to ensure that each of the

ecoregion that contains a site from the network. As

cells occupied by the species is contained within at least

earlier, this distance is coded to a gray level. Unlike

the Everglades-ness maps, however, darker areas one cluster. All clusters, when considered together as

multidimensional volumetric pixels, form a model of

represent areas that are poorly represented by the

the location and shape of the niche within data space.

existing network. Because this method quanti es

Selection of the radius (variance) for the clusters

coverage or presence of sites, sites will always sit

controls the resolution with which the niche envelope

within well-represented ecoregions, which will be col-

is de ned and serves to ll gaps within the Hutchin-

ored white.

sonian hypervolume.

Maps showing the geographic areas represented by

If the current geographic range data include some

each individual site can be generated, and importance

surrogate measures of tness under each combination

values can be calculated for each site based on the

of environmental conditions, this information can be

marginal representation it adds to the network.

attached to the clusters de ning the niche hypervo-

Quantifying the contribution of each site to network

lume. Hypervolume de nitions are used to make geo-

representation can minimize the impact of site elimi-

graphic range predictions for new areas by projecting

nation on representation. Finally, a network with a gi-

all map cells in the area into standardized environment

ven number of sites can be designed that is

space, then testing each cell to determine if it is within

theoretically optimal, having the highest possible rep-

one of the clusters that de ne the niche hypervolume.

resentation of environmental conditions on a map

The mean surrogate tness associated with the closest

(given that number of underlying ecoregions).

cluster centroid containing a cell is used to predict the

When submitted to a network analysis based on an

tness of the species within that cell location in the

ecoregionalization into the 2000 most different ecore-

new geographic range.

gions based on 25 environmental factors, the National

We de ned niche hypervolume models for loblolly

Science Foundation s Long-Term Ecological Research

pine, Pinus taeda L., and sugar maple, Acer saccharum

(LTER) study site network indicates that additional

Marsh., in terms of the 25 environmental condition

LTER sites in the Olympic peninsula and the Klamath,

gradients of climatic, physi



Contact this candidate