A nimal Learning & Behavior
A "random-walk" simulation model of
multiple-pattern learning in a radial-arm maze
IAN NEATHand E. J. CAPALDI
Purdue University, WestLafayette, Indiana
Wathen and Roberts (1994) reported rather surprising results of a radial maze study that on any in-
terpretation requires postulation of previously unsuspected high-level cognitive processes in rats. In
each of four arms of the eight-arm radial maze, a different serial pattern unfolded over trials; for ex-
ample, in one of the arms reward and nonreward alternated over successive trials. On each trial, rats
came to track successfully four different patterns simultaneously. The authors suggested that rats
tracked the pattern by using some form of trial-number strategy; that is, the trial number indicated
which arms contained the better rewards. This strategy could involve a hypothesis, considered unlikely
by some, that rats are capable of keeping track of as many as eight successive events-as, for exam-
ple, by counting. A simulation model that embodies a specific form of the trial-number hypothesis is
described here, and the results of the simulation correlate remarkably well with the observed data. In
addition, the model makes four separate predictions that are supported by Wathen and Roberts's data
and that seem beyond the scope ofother available theories.
Wathen and Roberts (1994) reported an experiment in these trials, the patterns were continued. Thus, the extrap-
olation trials for the single alternating pattern had no rein-
which rats were placed in an eight-arm radial maze and were
tested for their ability to learn multiple patterns of rein- forcement on Trial 9, reinforcement on Trial 10, and so on.
forcement simultaneously. For example, if Ann A of the The extrapolation trials for the 2R4N2R had four non-
maze were assigned a single alternation pattern, there would reinforced trials. These data are reproduced in Figure 1
be no food there on Trial 1, there would be food there on (closed circles).
Trial 2, there would be no food on Trial 3, there would be There was multiple-pattern tracking for Trials 1-8, where
food on Trial 4, and so on for eight trials. At the same time mean rank of entry into the arms closely followed whether
that the single alternation pattern was unfolding, three other the arm choice was reinforced or not. That is, when there
different patterns of reinforcement were unfolding simulta- were two pellets in an arm, it was entered earlier than the
neously in three other arms (see Table 1). Reinforced arms control arms, which had just one pellet; when an arm had
contained two 45-mg Noyes pellets, control arms contained no food, it was entered later than the control arms. The im-
one pellet, and nonreinforced arms contained no food. One portant point to remember is that all four patterns were being
pattern was single alternation (SA), beginning with nonre- tracked simultaneously. Remarkably, however, there ap-
inforcement; a second was double alternation (DA), begin- peared to be no evidence of pattern extrapolation on Trials
ning with reinforcement; a third had two nonreinforced tri- 9-12, even for the most simple pattern, single alternation.
als, four reinforced trials, and two nonreinforced trials
THEORIES OF MULTIPLE PATTERN
(2N4R2N); and the fourth had the opposite, two reinforced
LEARNING
trials, four nonreinforced trials, and two reinforced trials
(2R4N2R). The remaining four arms were control arms; pat-
terns were assigned randomly to the arms for each subject. Theories of multiple pattern learning may be divided into
After 14 days of habituation to the maze, during which two classes, the arm-pattern and the trial-number views. Both
time the animals learned to push a food cover to obtain one were originally designed to explain more simple pattern learn-
pellet and to climb over a barrier that blocked entry to each ing situations, but can be applied to multiple-pattern learn-
ing (see Wathen & Roberts, 1994). Within the arm-pattern
arm, the training sessions began. The main behavioral mea-
sure was the mean rank of entry into each of the eight arms class, one can further distinguish the sequential association
on each of the eight trials of each session. The main data of view and the rule learning view. Both of these may be
interest come from the 4 animals tested in Experiment 3, termed horizontal theories, because the focus is on the
which were also tested on four extrapolation trials. On changing patterns of reinforcement within a particular arm
over trials-or moving horizontally through Table 1.
Portions of this paper were presented at the Eighth Annual Indiana- The Sequential Association View of Capaldi
Purdue-Kentucky Conference on Animal Learning and Behavior, West
According to this view (Capaldi, 1992; Capaldi & Verry,
Lafayette, IN, March 1994. Correspondence may be addressed to
1981), animals learn sequences of serial associations as
I. Neath, 1364 Psychological Sciences Building, Purdue University,
chunks that are cued by prior events or memories. In an SA
West Lafayette, IN 47907-1364 (e-mail: *****@*****.******.***).
Copyright 1996 Psychonomic Society, Inc. 206
MULTIPLE-PATTERN LEARNING 207
c 4 4
III
III
:E
~ 3 3
2
2
1
6 8 8
10 0 2 4 12
0 2 4 12 6 10
Trial Trial
Data 2R4N2R
Data 2N4R2N
~
Model 2R4N2R
Model 2N4R2N -0-
:::E 3 3
2 2
1
a 8 12
10 2 10
4
8 6
0 2 12
2 10 12
0
Trial
Figure 1. Mean rank of entry as a function oftrial number for four patterns. The closed circles show data from
Wathen and Roberts (1994, Experiment 3); the open squares show the results ofthe simulation.
The Rule Learning View of Fountain and Hulse
pattern (NRNRNRNR), for example, memory of nonre-
ward can signal an upcoming rewarded trial. In a DA pat- According to this view-(Fountain, 1990; Fountain, Henne,
tern (RRNNRRNN), memory of a nonrewarded chunk & Hulse, 1984), animals learn hierarchical rules for gen-
can signal an upcoming rewarded chunk. In the case of erating patterns by using both working memory and ref-
multiple patterns, the animal would be required to keep erence memory. Working memory is used to associate an
track of the current trial so that the memory of reward (or event with a particular temporal/episodic context and is
nonreward) in a particular arm from the previous trial responsible for associating the sample stimulus with the
could signal the appropriate choice. trial. Reference memory, on the other hand, processes in-
NEATH AND CAPALDI
208
tions of numerical competence and counting ability, which
Table 1
would permit the trial number to be used as an important
Trial Number
cue (see, e.g., Capaldi, 1993; Capaldi & Miller, 1988).
5 6
3
I 4 7
Pattern 8
2
The trial-number hypothesis could be instantiated in
R N R N R
SA N R N
many ways. The random-walk simulation model, described
N R R N N
DA R R N
below, is one-perhaps the most simple-version of this
2N4R2N R R R R N N
N N
N
2R4N2R N N N R R
R R view. According to the model, the animals in the multiple-
I I
I I I I
Control I I
pattern learning experiment do not learn horizontal pat-
I I I
I I I
Control I I
terns; rather, they learn vertical patterns that indicate
I I I I I
I
Control I I
which arm contains the maximum reward on each of the
I I I I I
I I I
Control
eight trials.
Note-The patterns of reinforcement (R; two pellets) and nonrein-
forcement (N; zero pellets) for the four pattern arms are from Wathen
and Roberts (1994). Control patterns always had one pellet. SA, single THE RANDOM-WALK MODEL
alternation; DA, double alternation.
The random-walk model assumes that a rat can cor-
rectly identify the current trial number. For all trials dur-
ing the first session, all choices are random. As the animal
formation independently of the context and is responsible
experiences each reward level (zero, one, or two pellets),
for general rules and procedures. Formultiple-pattern learn-
it remembers which arm contained the largest reward, and
ing, reference memory is used to remember the overall se-
it associates that arm with the particular trial number. If
quence, such as A, then B, then C, and then D, whereas
working memory remembers the immediately preceding two or more arms have the same large reward, the animal
response to mark position within the pattern (Olton, randomly associates the trial number with one of the arms.'
Shapiro, & Hulse, 1984). As with the sequential associa- On subsequent sessions, for each trial, the animal recalls
which of the eight arms had the largest reward and enters
tion view, chunking can occur.
According to both of these horizontal views, if animals that arm first. The remaining seven arms are entered in
are tracking patterns, they should correctly extrapolate the random order; hence the name ofthe model. Note that pat-
pattern, particularly the simple single and double alterna- tern learning consists of learning which one arm had the
tion patterns, a prediction inconsistent with findings re- largest reward on each of eight trials and then randomly
ported by Wathen and Roberts (1994). For successful track- entering the remaining arms. The memory load is reduced
ing of multiple patterns, both views require that the animal from 64 items, according to the arm-pattern views, to 8
remember at least two sets of hierarchical rules or se- items. This reduction in memory load is achieved because
quential associations. For example, the rule-learning view of the assumption that the rat is capable of employing a
requires the animal to remember an 8 X 8 matrix of choices complex cognitive solution, keeping track of eight suc-
and to be able to identify correctly both the current trial cessive trials.
and the appropriate arm choice from 64 possibilities. Al- For all simulations, there were 28 sessions, as in Wa-
though chunking the patterns might reduce the memory then and Roberts's (1994) Experiment 3. For simplicity, it
load, the matrix of response choices would still exceed rea- was assumed that there was no forgetting of trial number
sonable estimates of memory capacity. Moreover, chunk- and that there was no forgetting ofthe largest reward mag-
ing would result in a matrix that is not square (e.g., a row nitude, although these could be added as parameters. The
of 8 for SA, a row of 4 for DA, and rows on for 2N4R2N simulation was run four times, and the data averaged to
and 2R4N2R), which would complicate the interaction represent running 4 rats in an experiment and averaging
between reference and working memory. In contrast to over each subject's data. The patterns and reinforcement
these horizontal theories are a class of vertical theories that contingencies were those used by Wathen and Roberts
emphasize the patterns of reinforcement in all arms on a (1994), and each virtual rat had a different random assign-
particular trial. These are "vertical" in the sense that the ment of patterns to arms, as in the original. Note that there
important comparison is the choices within a particular are no free parameters in the model: All ofthe parameters
trial, or moving vertically through Table 1. are specified in the study (e.g., number of arms, number
ofsubjects, number of trials, and so forth). The results are
The Trial-Number View shown in Figure 1 as open squares, along with the corre-
According to this view, rats keep track ofthe trial num- sponding data from Wathen and Roberts (1994, Figure 7).
ber, either by counting or by some other means (see, e.g., For Trials 1-8, the simulation is clearly replicating the
Holyoak & Patterson, 1981), and learn to expect reward in observed data. Particularly interesting are Trial 3 of the
certain arms on each trial. This view requires that the an- 2N4R2N and Trial 7 of the 2R4N2R patterns, because on
imal is able to identify correctly the current trial from eight these trials, there is only one arm with the maximum re-
possibilities. Given correct identification ofthe trial num- ward. In the simulation, all 4 simulated rats will choose to
ber, the animal then expects certain arms to contain rewards. enter this arm first, giving mean rank of entry of 1.0.2 The
There is no requirement that the animal relate patterns of rats in the Wathen and Roberts (1994) study showed an al-
reinforcement on trial n to the patterns of reinforcement on most identical level of performance, with mean rank of
trial n + 1. Data consistent with this view include sugges- entry of 1.5 and 1.4, respectively. Trials 1, 4, 5, and 8 had
MULTIPLE-PATTERN LEARNING 209
two arms that had the maximum reward. The average rank here, both horizontal views previously mentioned have the
of entry, in the simulation, for the arms with two pellets on same requirement in order to explain the Wathen and
these trials was 2.0. On Trials 2 and 6, there were three Roberts (1994) data. The difference between the horizontal
arms that had the maximum reward. The average rank of and vertical views is that any horizontal view must neces-
entry for these three arms was 2.5. The simulation model sarily make an additional assumption of accurate recollec-
predicts that the more arms there are with the maximum tion of eight times more information: In addition to tracking
reward, the later the mean rank of entry, exactly the pat- trial number, the animal must also track the information
tern observable in the data. The reason, according to the contained in the columns of Table I. Thus, the trial num-
model, is that the rats randomly associate one of the arms ber assumption is indispensable for any model.
with the maximum reward as the arm to enter first. The The random-walk model is easily tested, in principle,
other arms with a maximum reward will be entered later because it makes at least four strong predictions that are
(on the average, fifth) and this will inflate the mean rank supported by the Wathen and Roberts (1994) data. First,
of entry. as previously indicated, the model predicts no extrapola-
The extrapolation trials were simulated by having the tion, regardless of the pattern, unless the subject believes
model randomly pick each arm to enter. Consider the fol- that the first extrapolation trial is actually Trial I again.
lowing analogy. A human subject is presented with eight There is no extrapolation because the subject is using the
items in random order and is asked to recall all eight. On trial number as a cue, and when the trial number is invalid,
an extrapolation trial, the subject is now asked to recall the there can be no memory to guide the choice of an arm.
ninth item. Ifpressed for a response, the subject would have Second, the model predicts highly stereotyped pattern of
to guess (randomly pick a response), because the question arm choice. Because the animals are learning which arm
does not make sense. The simulation assumes that the an- contains the maximum reward on that trial, they will al-
imals are recalling the trial number, but when this number ways enter that arm very early, producing the very strong
preferences for the initial arm choice of a trial observed by
is not valid, the animals randomly pick which arms to enter.
Wathen and Roberts. Third, the model predicts that mas-
Again, the simulation results mimic the observed data.
tery of a horizontal pattern is unrelated to the regularity of
The correlations between the observed and predicted
the pattern. For example, a random pattern could be learned
data were .958 for SA, .868 for DA, .855 for 2N4R2N, and
better than a single alternation pattern because it is the ver-
.917 for 2R4N2R for the first eight trials, and they were
tical patterns that are more important than the horizontal
.922, .858, .822, and .916 when the extrapolation trials
patterns. Thus, in the Wathen and Roberts study, the rats
were also included. The fits are by no means perfect, but
"learned" the irregular 2N4R2N pattern as well as they
they are certainly suggestive. A vertical model with no free
learned the more regular DA pattern. Fourth, the model
parameters, which assumes only that the rat can recall the
predicts less irregularity in performance when, for all tri-
arm with the largest reward on each trial, produces results
als, there is an equal number of arms that contain the max-
that closely mimic horizontal multiple-pattern learning. In
addition, it predicts that there will be no extrapolation of imum reward. For example, the two earliest ranks of entry
were observed when there was only one arm that contained
the pattern, even for the most simple and regular of pat-
terns, because horizontal pattern learning did not occur. the maximum reward. When the maximum reward was
Again, this is what was observed. available in three arms, the mean rank of entry was later.
It is important to note that the random-walk model is
IMPLICATIONS AND PREDICTIONS not applicable when single patterns are to be tracked: The
model applies only when the task places a high load on the
In 1972, Underwood posed a question to researchers cognitive abilities of the organism. According to the model,
studying human behavior: "Are we overloading memory?" when this situation occurs, the animal copes by reverting
to a mnemonically simpler but cognitively complex strat-
The same question may be asked of multiple-pattern
learning theorists. Both horizontal views previously dis- egy of associating trial number or position with the max-
imum reward available on that trial. Although most of the
cussed can account for pattern learning when only a single
pattern is tracked. However, when multiple patterns are decisions (56 out of 64) will be random, the remaining 8
tracked, these views require the animal to remember a sub- nonrandom decisions will give rise to pattern learning.
stantial amount of information; in the case of the Wathen One interesting question to explore concerns how high
the cognitive load must be before the animal uses the
and Roberts (1994) study, an 8 X 8 matrix. Moreover,
random-walk strategy. Cognitive load, however, should
both horizontal views predict accurate extrapolation, par-
ticularly for the relatively simple single and double alter- not necessarily be interpreted as a capacity limitation based
nation patterns. No evidence of extrapolation was found. on the number of items. In the human literature, for exam-
The vertical view offered here, the random-walk simu- ple, many theories of working memory are moving away
from item- or time-based limits of capacity and instead
lation model, predicts a failure to extrapolate. This model
assumes that the animal has to remember only 8, rather focus on resource or discrimination limits (e.g., Nairne, in
than 64, items. Interestingly, even if one considers it ques- press). It is likely that capacity limitations in the rat, par-
tionable to assume that a rat can correctly identify the cur- ticularly in situations of high cognitive load, do not depend
rent trial number, note that like the vertical view presented on the absolute number of items or the absolute passage of
210 NEATH AND CAPALDI
time, but rather on the animal's ability to discriminate or in rat serial-pattern tracking. Journal of Experimental Psychology:
differentiate between two or more alternatives on the basis Animal Behavior Processes, 16,96-105.
FOUNTAIN, S. B., HENNE,D. R., & HULSE, S. H. (1984). Phrasing cues
of cues available at retrieval (cf. Neath & Knoedler, 1994).
and hierarchical organization in serial pattern learning by rats. Journal
Historically, two approaches to serial learning-in both of Experimental Psychology: Animal Behavior Processes, 10, 30-45.
humans and animals-have been dominant competitors HOLYOAK, K. 1., & PATTERSON, K. K. (1981). A positional discrim-
(cf. Crowder, 1976). The item-item view holds that prior inability model of linear-order judgments. Journal of Experimental
Psychology: Human Perception & Performance, 7,1283-1302.
items in a series provide cues for later items in the series.
NAIRNE, J. S. (in press). Short-term/working memory. In E. L. Bjork &
The position-item view holds that cues associated with ei-
R. A. Bjork (Eds.), Handbook ofperception and cognition: Vol. 10.
ther the absolute or relative position of an item in a series Memory. New York: Academic Press.
provide cues for items in the series. The trial-number view NEATH, I., & KNOEDLER, A. J. (1994). Distinctiveness and serial posi-
advanced here is obviously a version of the position-item tion effects in recognition and sentence processing. Journal of Mem-
ory & Language, 33, 776-795.
view, but it differs from previous versions: The cue for an
OLTON, D. S., SHAPIRO, M. L., & HULSE, S. H. (1984). Working mem-
item identified here is not position in a series of items but ory and serial patterns. In H. L. Roitblat, T. G. Bever, & H. S. Terrace
rather position in a series of series. (Eds.), Animal cognition (pp. 171-182). Hillsdale, NJ: Erlbaum.
UNDERWOOD, B. 1. (1972). Are we overloading memory? In A. W.
REFERENCES Melton & E. Martin (Eds.), Coding processes in human memory
(pp. 1-23). New York: Wiley.
CAPALDI, E. J. (1992). Levels of organized behavior in rats. In W. K. WATHEN, C. N., & ROBERTS, W. A. (1994). Multiple-pattern learning by
Honig & 1. G. Fetterman (Eds.), Cognitive aspects ofstimulus control rats on an eight-arm radial maze. Animal Learning & Behavior, 22,
(pp. 385-404). Hillsdale, NJ: Erlbaum. . 155-164.
CAPALDI, E. J. (1993). Animal number abilities: Implications for a hier-
NOTES
archical approach to instrumental learning. In S. Boysen & E. 1. Ca-
paldi (Eds.), The development of numerical competence (pp. 191-
209). Hillsdale, NJ: Erlbaum. I. This need not imply that a particular rat necessarily remembers the
CAPALDI, E. J., & MILLER, D. J. (1988). Counting in rats: Its functional two (or more) alternatives and picks one at random. Rather, it means only
significance and the independent cognitive processes that constitute it. that the basis of the choice is random. For example, if both Arm D and
Arm G contain the maximum reward for a particular trial, n, Rat A might
Journal ofExperimental Psychology: Animal Behavior Processes, 14,
3-17. enter Arm D first and associate Arm D with trial n whereas Rat B might
CAPALDI, E. 1., & VERRY, D. R. (1981). Serial order anticipation learn- enter Arm G first and associate Arm G with trial n.
ing in rats: Memory for multiple hedonic events and their order. Ani- 2. The actual values were 1.13 and 1.12, respectively, because on the
mal Learning & Behavior, 9, 441-453. first trial all arms are entered randomly.
CROWDER, R. G. (1976). Principles of learning and memory. Hillsdale,
NJ: Erlbaum. (Manuscript received February 10, 1995;
FOUNTAIN, S. B. (1990). Rule abstraction, item memory, and chunking revision accepted for publication May 13, 1995.)
Copyright 1996 Psychonomic Society, Inc. 206