{"title": "A Cortico-Cerebellar Model that Learns to Generate Distributed Motor Commands to Control a Kinematic Arm", "book": "Advances in Neural Information Processing Systems", "page_first": 611, "page_last": 618, "abstract": null, "full_text": "A Cortico-Cerebellar Model that Learns to \nGenerate Distributed Motor Commands to \n\nControl a Kinematic Arm \n\nN.E. Berthier S.P. Singh A.G. Barto \n\nDepartment of Computer Science \n\nUniversity of Massachusetts \n\nAmherst, MA 01002 \n\n.T.C. Honk \n\nDepartment of Physiology \n\nNorthwestern University Medical School \n\nChicago, IL 60611 \n\nAbstract \n\nA neurophysiologically-based model is presented that controls a simulated \nkinematic arm during goal-directed reaches. The network generates a \nquasi-feedforward motor command that is learned using training signals \ngenerated by corrective movements. For each target, the network selects \nand sets the output of a subset of pattern generators. During the move(cid:173)\nment, feedback from proprioceptors turns off the pattern generators. The \ntask facing individual pattern generators is to recognize when the arm \nreaches the target and to turn off. A distributed representation of the mo(cid:173)\ntor command that resembles population vectors seen in vivo was produced \nnaturally by these simulations. \n\n1 \n\nINTRODUCTION \n\nWe have recently begun to explore the properties of sensorimotor networks with \narchitectures inspired by the anatomy and physiology of the cerebellum and its in(cid:173)\nterconnections with the red nucleus and the motor cortex (Houk 1989; Houk et al.. \n611 \n\n\f612 \n\nBerthier, Singh, Barto, and Houk \n\n1990). It is widely accepted that these brain regions are important in the control \nof limb movements (Kuypers, 1981; Ito, 1984), although relatively little attention \nhas been devoted to probing how the different regions might function together in \na cooperative manner. Starting from a foundation of known anatomical circuitry \nand the results of microelectrode recordings from neurons in these circuits, we pro(cid:173)\nposed the concept of rubrocerebellar and corticocerebellar information processing \nmodules that are arranged in parasagittal arrays and function as adjustable pattern \ngenerators (APGs) capable of the storage, recall and execution of motor programs. \n\nThe aim of the present paper is to extend the APG Model to a multiple degree(cid:173)\nof-freedom task and to investigate how the motor representation developed by the \nmodel compares to the population vector representations seen by Georgopoulos \nand coworkers (e.g., Georopoulos, 1988). A complete description of the model and \nsimulations reported here is contained in Berthier et al. (1991). \n\n2 THE APG ARRAY MODEL \n\nAs shown in Figure 1 the model has three parts: a neural network that generates \ncontrol signals, a muscle model that controls joint angle, and a planar, kinematic \narm. The control network is an array of APGs that generate signals that are \nfed to the limb musculature. Because here we are interested in the basic issue of \nhow a collection of APGs might cooperatively control multiple degree-of-freedom \nmovements, we use a very simplified model of the limb that ignores dynamics. The \nmuscles convert APG activity to changes in muscle length, which determine the \nchanges in the joint angles. Activation of an APG causes movement of the arm in \na direction in joint-angle space that is specific to that APG 1 , and the magnitude \nof an APG's activity determines the velocity of that movement. The simultaneous \nactivation of selected APGs determines the arm trajectory as the superposition of \nthese movements. A learning rule, based on long-term depression (e.g., Ito, 1984), \nadjusts the subsets of APGs that are selected as well as characteristics of their \nactivity in order to achieve desired movements. \n\nEach APG consists of a positive feedback loop and a set of Purkinje cells (PCs). \nThe positive feedback loop is a highly simplified model of a component of a complex \ncerebrocerebellar recurrent network. In the simplified model simulated here, each \nAPG has its own feedback loop, and the loops associated with different APGs do \nnot interact. When triggered by sufficiently strong activation, the neurons in these \nloops fire repetitively in a self-sustaining manner. An APG's motor command is \ngenerated through the action of its PCs which inhibit and modulate the buildup of \nactivity in the feedback loop. The activity of loop cells is conveyed to spinal motor \nareas by rubrospinal fibers. PCs receive information that specifies and constrains \nthe desired movements via parallel fibers. \n\nWe hypothesize that the response of PCs to particular parallel fiber inputs is adap(cid:173)\ntively adjusted through the influence of climbing fibers that respond to corrective \nmovements (Houk & Barto, 1991). The APG array model assumes that climbing \nfibers and PCs are aligned in a way that climbing fibers provide specialized infor-\n\nITo simplify these initial simulations we ignore changes in muscle moment arms with \n\nposture of the arm. \n\n\fA Cortico-Cerebellar Model that Learns to Generate Distributed Motor Commands \n\n613 \n\nNetwork \n\n........................... , \u00b7 \u00b7 \u00b7 \nAPG Modules l \n\nMuscles \n\n1 \n\nm \n\nM \n\n\u00b7 \u00b7 \u00b7 \u00b7 \n\n................................... .1 \n\nT \u2022 \n\nT \u2022 \n\nFigure 1: APG Control of Joint Angles. A collection of of APGs (adjustable pattern \ngenerators) is connected to a simulated two degree-of-freedom, kinematic, planar \narm with antagonistic muscles at each joint. The task is to move the arm in the \nplane from a central starting location to one of eight symmetrically placed targets. \nActivation of an APG causes a movement of the arm that is specific to that APG, \nand the magnitude of an APG's activity determines the velocity of that movement. \nThe simultaneous activation of selected APGs determines the arm trajectory as a \nsuperposition of these movements. \n\nmation to PCs. Gellman et al. (1985) showed that proprioceptive climbing fibers \nare inhibited during planned movements, but the data of Gilbert and Thach (1977) \nsuggest that they fire during corrective movements. In the present simulations, we \nassume that corrective movements are made when a movement fails to reach the \ntarget. These corrective movements stimulate proprioceptive climbing fibers which \nprovides information to higher centers about the direction of the corrective move(cid:173)\nment. More detailed descriptions of APGs and relevant anatomy and physiology \ncan be found in Houk (1989), Houk et al. (1990), and Berthier et al. (1991). \n\nThe generation of motor commands occurs in three phases. In the first phase, we \nassume that all positive feedback loops are off, and inputs provided by teleceptive \nand proprioceptive parallel fibers and basket cells determine the outputs of the PCs. \nWe call this first phase selection. We assume that noise is present during the selec(cid:173)\ntion process so that individual PCs are turned off (Le., selected) probabilistic ally. \nTo begin the second phase, called the execution phase, loop activity is triggered by \ncortical activity. Once triggered, loop activity is self-sustaining because the loop \ncells have reciprocal positive connections. The triggering of loop activity causes the \nmotor command to be \"read out.\" The states of the PCs in the selection phase \ndetermine the speed and direction of the arm movement. As the movement is be(cid:173)\ning performed, proprioceptive feedback and efference copy gradually depolarize the \nPCs. When a large proportion of the PCs are depolarized, PC inhibition reaches a \ncritical value and terminates loop activity. In the third phase, the correction phase, \ncorrective movements trigger climbing fiber activity that alters parallel fiber-PC \nconnection weights. \n\n\fB \n\n614 \n\nBerthier, Singh, Barto, and Houk \n\nA \n\n\u2022 \n\nI \n\n\u2022\n\nI . I I \n\n\u2022 \u2022 ' \n\nI . I \n\n\u2022 \n\nI,' \n\u2022 : . ~ . . . . . \n\n11-: .... \nw..J \n\n[J ... \n, . \n\" . \n::: \n'. '. II .,: \u2022 I.:' \n. '. I.. .::. \n. ..' \nn,.\u00b7\u00b7 .. ....... ,:1. \n.~ .1-... tI , \n.\". ~\" \n..\u2022 r \nL.!I \n:. \n. \n~ \n6 \n\n.. \u2022 . \n. ' \n~ . \nc1: \n\n~ \n\u2022 \n\nI\" \n' \n\nI' \u2022 \n\n.1 \u2022 \n\n' . \n\n\u2022 \n\n\u2022 \n\nrI \n\n. . . u \n\n, ~ \u2022 \n\n~ \nLJ \n\n\u2022 \n\n.. \n\n' ... , \n. \n'0 \n\nII \n\n\u2022 \n\nFigure 2: A. Movement Trajectories After Training. The starting point for each \nmovement is the center of the workspace, and the target location is the center of \nthe open square. The position of the arm at each time step is shown as a dot. \nThree movements are shown to each target. B. APG selection. APG selection for \nmovements to a given target is illustrated by a vector plot at the position of the \ntarget. An individual APG is represented by a vector, the direction of which is \nequal to the direction of movement caused by that APG in Cartesian space. The \nvector length is proportional to output of the Purkinje cells during the selection \nphase. The arrow points in the direction of the vector sum. \n\n3 SIMULATIONS \n\nWe trained the APG model to control a two degree-of-freedom, kinematic, planar \narm. The task was similar to Georgopoulos (1988) and required APGs to move the \narm from a central starting point to one of eight radially symmetric, equidistant \ntargets. Each simulated trial started by placing the endpoint of the arm in the cen(cid:173)\ntral starting location. The selection, execution, and correction phases of operation \nwere then simulated. The task facing each of the selected APGs was to turn off at \nthe proper time so that the movement stopped at the target. \n\nSimulations showed that the model could learn to control movements to the eight \ntargets. Training typically required about 700 trials per target until the arm end(cid:173)\npoint was consistently moved to within 1 em of the target. Figure 2 shows sample \ntrajectories and population vectors of APG activity. Performance never resulted \nin precise movements due to the probabilistic nature of selection. Movement tra(cid:173)\njectories tended to follow straight lines in joint-angle space and were thus slightly \ncurved lines in the workspace. About half of the APGs in the model were used \nto move to an individual target with population vectors similar to those seen by \nGeorgopoulos (1988). The number of APGs used for each target was dependent \non the sharpness of the climbing fiber receptive fields, with cardioid shaped recep(cid:173)\ntive fields in joint-angle space giving population vectors that most resembled those \nexperimentally observed. \n\n\fA Carrico-Cerebellar Model that Learns to Generate Distributed Motor Commands \n\n615 \n\n4 ANALYSIS \n\nIn order to understand how the model worked we undertook a theoretical analysis of \nits simulated behavior. Analysis indicated that the expected trajectory of a move(cid:173)\nment was a straight line in joint-angle space from the starting position to the target. \nThis is a special case of a mathematical result by Mussa-Ivaldi (1988). Because se(cid:173)\nlection is probabilistic in the APG Array Model, trajectories in the workspace varied \nfrom the expected trajectory. In these cases, trajectories were piecewise linear be(cid:173)\ncause of the asynchronous termination of APG activity. Because of the Law of \nLarge Numbers, the more PCs in each APG, the more closely the movement will \nresemble the expected movement. \n\nThe expected population of vectors of APG activity can be shown to be cosine(cid:173)\nshaped in joint-angle space. That is, the length of the vector representing the \nactivity of APG m is proportional to the cosine of the angle between the direction \nof action of APG m and the direction of the target in joint-angle space. The shape \nof the population vectors in Cartesian space is dependent on the Jacobian of the \narm, which is a function of the arm posture. \n\nThe manner in which the outputs of PCs were set during selection leads to scaling of \nmovement velocity with target distance. For any given movement direction, targets \nthat are farther from the starting location lead to more rapid movements than closer \ntargets. \n\nUpdating network weights based on the expected corrective movement will, in some \ncases, result in changing the weights in a way that they converge to the correct \nvalues. However, in other cases inappropriate changes are made. In the current \nsimulations, we could largely avoid this problem by selecting parameter and initial \nweight values so that movements were initially small in amplitude. Random initial(cid:173)\nization of the weight values sometimes led to instances from which the learning rule \ncould not recover. \n\n5 DISCUSSION \n\nIn general, the present implementation of the modelled to adequate control of the \nkinematic arm and mimicked the general output of nervous system seen in actual \nexperiments. The network implemented a spatial to temporal transformation that \ntransformed a target location into a time varying motor command. The model \nnaturally generated population vectors that were similar to those seen in vivo. Fur(cid:173)\nther research is needed to improve the model's robustness and to extend it to more \nrealistic control of a dynamical limb. \n\nIn the APG array model, APGs control arm movement in parallel so that the activ(cid:173)\nity of all the modules taken together forms a distributed representation. The APG \narray executes a distributed motor program because it produces a spatiotemporal \npattern of activity in the cerebrocerebellar recurrent network that is transmitted to \nthe spinal cord to comprise a distributed motor command. \n\n\f616 \n\nBerthier, Singh, Barto, and Houk \n\n5.1 PARAMETRIZED MOTOR PROGRAMS \n\nCertain features of the APG array model relate well to the ideas about parameter(cid:173)\nized motor programs discussed by Keele (1973), Schmidt (1988), and Adams (1971, \n1977). The selection phase of the APG array model provides a feasible neuronal \nmechanism for preparing a parameterized motor program in advance of movement. \nThe execution phase is also consistent with the open-loop ideas associated with \nmotor programming concepts, except that, like Adams (1977), we explain the ter(cid:173)\nmination of the execution phase as being a consequence of proprioceptive feedback \nand efference copy. \n\nIn the APG array model, the counterpart of a generalized motor program is a \nset of parallel fiber weights for proprioceptive, efference copy, and target inputs. \nGiven these weights, a particular constellation of parallel fiber inputs signifies that \nthe desired endpoint of a movement is about to be reached, causing PCs to become \ndepolarized. Once a set of parallel fiber weights corresponding to a desired endpoint \nis learned, the neuronal architecture and neurodynamics of the cerebellar network \nfunctions in a manner that parameterizes the motor program. \n\nMovement velocity is parameterized in the selection phase of the model's operation. \nThe velocity that is selected is automatically scaled so that velocity increases as the \namplitude of the movement increases. While this type of scaling is often observed \nin motor performance studies, velocity can also be varied in an independent man(cid:173)\nner where velocity scaling can be applied simultaneously to all elements of a motor \nprogram to slow down or speed up the entire movement. Although we have not ad(cid:173)\ndressed this issue in the present report, simulation of velocity scaling under control \nof a neuromodulator can naturally be accomplished in the APG array model. \n\nMovements terminate when the endpoint is recognized by PCs so that movement \nduration is dependent on the course of the movement instead of being determined by \nsome internal clock because. Movement amplitude is parameterized by the weights \nof the target inputs, with smaller weights corresponding to larger amplitude move(cid:173)\nments. \n\n5.2 CORRECTIVE MOVEMENTS \n\nWe assume that the training information conveyed to the APGs is the result of crude \ncorrective movements stimulating proprioceptive receptors. This sensory informa(cid:173)\ntion is conveyed to the cerebellum by climbing fibers. Learning in the APG array \nmodel therefore requires the existence of a low-level system capable of generating \nmovements to spatial targets with at least a ballpark level of accuracy. Lesion (Yu \net al., 1980) and developmental studies (von Hofsten, 1982) support the existence \nof a low-level system. Other evidence indicates that when limb movements are not \nproceeding accurately toward their intended targets, corrective components of the \nmovements are generated by an unconscious, automatic control system (Goodale et \naI., 1986). \n\nWe assume that collaterals from the corticospinal and rubrospinal system that con(cid:173)\nvey the motor commands to the spinal cord gate off sensory transmission through \nthe proprioceptive climbing fiber pathway, thus preventing sensory responses to the \ninitial limb movement. As the initial movement proceeds, the low-level system re-\n\n\fA Corrico-Cerebellar Model that Learns to Generate Distributed Motor Commands \n\n617 \n\nceives proprioceptive feedback from the limb and feedforward information about \ntarget location from the gaze control system. The latter information is updated as \na consequence of corrective eye movements that typically occur after an initial gaze \nshift toward a visual target. Updated gaze information causes the spinal proces(cid:173)\nsor to generate a corrective component that is superimposed on the original motor \ncommand (Gielen & van Gisbergen, 1990; Flash & Henis, 1991). Since climbing \nfiber pathways would not be gated off by this low-level corrective process, climbing \nfibers should fire to indicate the direction of the corrective movement. \n\nWe assume that the network by which climbing fiber activity is generated is specif(cid:173)\nically wired to provide appropriate training information to the APGs (Houk & \nBarto, 1991). The training signal provided by a climbing fiber is specialized for the \nrecipient APG in that it provides directional information in joint-angle space that is \nrelative to the direction in which that APG moves the arm. The fact that training \ninformation is provided in terms of joint-angle space greatly simplifies the problem \nof providing errors in the correct system of reference. For example, if the network \nused visual error information, the error information would have to be transformed \nto joint errors. \n\nThe specialized training signals provided by the climbing fibers are determined by \nthe structure of the ascending network conveying proprioceptive information. This \nascending network has the same structure-but works in the opposite direction-as \nthe network by which the APG array influences joint movement. This is reminis(cid:173)\ncent of the error backpropagation algorithm (e.g., Rumelhart et al., 1986, Parker, \n1985) where the forward and backward passes through the network in the back(cid:173)\npropagation algorithm are accomplished by the descending and ascending networks \nof the APG Array Model. This use of the ascending network to transform errors in \nthe workspace to errors that are relative to a particular APG's direction of action \nis closely related to the use of error backpropagation for \"learning with a distal \nteacher\" as suggested by Jordan and Rumelhart (1991). \n\nHouk and Barto (1991) suggested that the alignment of the ascending and de(cid:173)\nscending networks might come about through trophic mechanisms stimulated by \nuse-dependent alterations in synaptic efficacy. In the context of the present model, \nthis hypothesis implies that the ascending network to the inferior olive, is estab(cid:173)\nlished first, and that the descending network by which APGs influence motoneurons \nchanges. We have not yet simulated this mechanism to see if it could actually gen(cid:173)\nerate the kind of alignment we assume in the present model. \n\nAcknowledgements \n\nThis research was supported by ONR N00014-88-K-0339, NIMH Center Grant P50 \nMH48185, and a grant from the McDonnell-Pew Foundation for Cognitive Neuro(cid:173)\nscience supported by the James S. McDonnell Foundation and the Pew Charitable \nTrusts. \n\nReferences \n\nAc!ams JA (1971) A closed-loop theory of motor learning. J Motor Beh 3: 111-149 \nAdams J A (1977) Feedback theory of how joint receptors regulate the timing and \n\npositioning of a limb. Psychol Rev 84: 504-523 \n\n\f618 \n\nBerthier, Singh, Barto, and Houk \n\nBerthier NE Singh SP Barto AG Houk JC (1991) Distributed representation of \nlimb motor programs in arrays of adjustable pattern generators. NPB Technical \nReport 3, Institute for Neuroscience, Northwestern University, Chicago IL \n\nFlash T Henis E (1991) Arm trajectory modifications during reaching towards \n\nvisual targets. J Cognitive Neurosci 3:220-230 \n\nGellman R Gibson AR Houk JC (1985) Inferior olivary neurons in the awake cat: \n\nDetection of contact and passive body displacement. J N europhys 54:40-60. \n\nGeorgopoulos A (1988) Neural integration of movement: role of motor cortex in \n\nreaching. FASEB Journal 2:2849-2857. \n\nGielen CCAM Gisbergen van JAM (1990) The visual guidance of saccades and \n\nfast aiming movements. News in Physiol Sci 5: 58-63 \n\nGilbert PFC Thach WT (1977) Purkinje cell activity during motor learning. Brain \n\nRes 128:309-328. \n\nGoodale MA Pelisson D Prablanc C (1986) Large adjustments in visually guided \n\nreaching do not depend on vision of the hand or perception of target displace(cid:173)\nment. Nature 320: 748-750 \n\nHofsten von C (1982) Eye-hand coordination in the newborn. Dev Psycho I 18: \n\n450-461 \n\nHouk JC (1989) Cooperative control of limb movements by the motor cortex, \nbrainstem and cerebellum. In: Cotterill RMJ (ed) Models of Brain Function. \nCambridge Univ Press Cambridge UK, 309-325 \n\nHouk JC Barto AG (1991) Distributed sensorimotor learning. NPB Technical \n\nReport 1, Institute for Neuroscience, Northwestern University, Chicago IL \n\nHouk JC Singh SP Fisher C Barto AG (1990) An adaptive sensorimotor network \n\ninspired by the anatomy and physiology of the cerebellum. In: Miller WT Sut(cid:173)\nton RS Werbos PJ (eds) Neural Networks for Control. MIT Press Cambridge, \nMA 301-348 \n\nIto M (1984) The Cerebellum and Neural Control. Raven Press New York \nIto M (1989) Long-term depression. Annual review of Neuroscience 12: 85-102 \nJordan MI Rumelhart DE (1991) Forward models: Supervised learning with a \n\ndistal teacher. Occasional Paper #40 MIT Center for Cognitive Science \n\nKeele SW (1973) Attention and Human Performance. Goodyear Pacific Palisades, \n\nCalifornia \n\nKuypers HGJM (1981) Anatomy of the descending pathways. In: Brooks VB (ed) \nHandbook of Physiology Section I Volume II Part 1. American Physiological \nSociety Bethesda MD 597-666 \n\nMussa-Ivaldi FA (1988) Do neurons in the motor cortex encode movement direc(cid:173)\n\ntion? An alternative hypothesis. Neurosci Lett 91:106-111 \n\nParker DB (1985) Learning-Logic. Technical Report TR-47, Massachusetts Insti(cid:173)\n\ntute of Technology Cambridge MA \n\nRumelhart DE Hinton GE Williams RJ (1986) Learning internal representations \n\nby error propagation. In: Rumelhart DE McClelland JL (eds) Parallel Dis(cid:173)\ntributed Processing. Explorations in the Microstructure of Cognition, Vol. 1: \nFoundations. Bradford Books/MIT Press Cambridge MA \n\nSchmidt RA (1988) Motor Control and Motor Learning. Human Kinetics Cham(cid:173)\n\npaign, Illinois \n\n\f", "award": [], "sourceid": 532, "authors": [{"given_name": "N.", "family_name": "Berthier", "institution": null}, {"given_name": "S. P.", "family_name": "Singh", "institution": null}, {"given_name": "A.", "family_name": "Barto", "institution": null}, {"given_name": "J. C.", "family_name": "Houk", "institution": null}]}