{"title": "Tonal Music as a Componential Code: Learning Temporal Relationships Between and Within Pitch and Timing Components", "book": "Advances in Neural Information Processing Systems", "page_first": 1085, "page_last": 1092, "abstract": null, "full_text": "Tonal Music as a Componential Code: \n\nLearning Temporal Relationships Between and \n\nWithin Pitch and Timing Components \n\nCatherine Stevens \n\nDepartment of Psychology \nUniversity of Queensland \n\nQLD 4072 Australia \n\nkates@psych.psy.uq.oz.au \n\nJanet Wiles \n\nDepts of Psychology & Computer Science \n\nUniversity of Queensland \nQLD 4072 Australia \njanetw@cs.uq.oz.au \n\nAbstract \n\nThis study explores the extent to which a network that learns the \ntemporal relationships within and between the component features of \nWestern tonal music can account for music theoretic and psychological \nphenomena such as the tonal hierarchy and rhythmic expectancies. \nPredicted and generated sequences were recorded as the representation of \na 153-note waltz melody was learnt by a predictive, recurrent network. \nThe network learned transitions and relations between and within pitch \nand timing components: accent and duration values interacted in the \ndevelopment of rhythmic and metric structures and, with training, the \nnetwork developed chordal expectancies in response to the activation of \nindividual tones. Analysis of the hidden unit representation revealed \nthat musical sequences are represented as transitions between states in \nhidden unit space. \n\n1 \n\nINTRODUCTION \n\nThe fundamental features of music, derivable from frequency, time and amplitude \ndimensions of the physical signal, can be described in terms of two systems - pitch and \ntiming. The two systems are frequently disjoined and modeled independently of one \nanother (e.g. Bharucha & Todd, 1989; Rosenthal, 1992). However, psychological \nevidence suggests that pitch and timing factors interact (Jones, 1992; Monahan, Kendall \n\n1085 \n\n\f1086 \n\nStevens and Wiles \n\n& Carterette, 1987). The pitch and timing components can be further divided into tone, \noctave, duration and accent which can be regarded as a quasi-componential code. The \nimportant features of a componential code are that each component feature can be viewed \nas systematic in its own right (Fodor & Pylyshyn, 1988). The significance of \ncomponential codes for learning devices lies in the productivity of the system - a small \n(polynomial) number of training examples can generalise to an exponential test set \n(Brousse & Smolensky, 1989; Phillips & Wiles, 1993). We call music a quasi(cid:173)\ncomponential, code as there are significant interactions between, as well as within, the \ncomponent features (we adopt this term from its use in describing other cognitive \nphenomena, such as reading, Plaut & McClelland, 1993). \n\nConnectionist models have been developed to investigate various aspects of musical \nbehaviour including composition (Hild, Feulner & Menzel, 1992), performance (Sayegh, \n1989), and perception (Bharucha & Todd, 1989). The models have had success in \ngenerating novel sequences (Mozer, 1991), or developing properties characteristic of a \nlistener, such as tonal expectancies (Todd, 1988), or reflecting properties characteristic of \nmusical structure, such as hierarchical organisation of notes, chords and keys (Bharucha, \n1992). Clearly, the models have been designed with a specific application in mind and, \nalthough some attention has been given to the representation of musical information \n(e.g. Mozer, 1991; Bharucha, 1992), the models rarely explain the way in which musical \nrepresentations are constructed and learned. These models typically process notes which \nvary in pitch but are of constant duration and, as music is inherently temporal, the \ntemporal properties of music must be reflected in both representation and processing. \nThere is an assumption implicit in cognitive modeling in the in/ormation processing \nframework that the representations used in any cognitive process are specified a priori \n(often assumed to be the output of a perceptual process). In the neural network \nframework, this view of representation has been challenged by the specification of a dual \nmechanism which is capable of learning representations and the processes which act upon \nthem. Since the systematic properties of music are inherent in Western tonal music as \nan environment, they must also be reflected in its representation in, and processes of, a \ncognitive system. Neural networks provide a mechanism for learning such \nrepresentations and processes, particularly with respect to temporal effects. \n\nIn this paper we study how representations can be learned in the domain of Western tonal \nmusic. Specifically, we use a recurrent network trained on a musical prediction task to \nconstruct representations of context in a musical sequence (a well-known waltz), and then \ntest the extent to which the learned representation can account for the classic phenomena \nof music cognition. We see the representation construction as one aspect of learning and \nmemory in music cognition, and anticipate that additional mechanisms of music \ncognition would involve memory processes that utilise these representations (e.g. \nStevens & Latimer, 1992), although they are beyond the scope of the present paper. For \nexample, Mozer (1992) discusses the development of higher-order, global representations \nat the level of relations between phrases in musical patterns. By contrast, the present \nstudy focusses on the development of representations at the level of relations between \nindividual musical events. We expect that the additional mechanisms would, in part, \ndevelop from the behavioural aspects of music cognition that are made explicit at this \nearly, representation construction stage. \n\n\fTonal Music as a Componential Code \n\n1087 \n\n2 METHOD & RESULTS \n\nA simple recurrent network (Elman. 1989), consisting of 25 input and output units, and \n20 hidden and context units. was trained according to Elman's prediction paradigm and \nused Backprop Through Time (BPlT) for one time step. The training data comprised the \nnrst 2 sections of The Blue Danube by Johann Strauss wherein each training pattern \nrepresented one event, or note. in the piece, coded as components of tone (12 values), rest \n(1 value), octave (3 values), duration (6 values) and accent (3 values). \n\nIn the early stages of training, the network learned to predict the prior probability of \nevents (see Figure 1). This type of information could be encoded in the bias of the \noutput units alone. as it is independent of temporal changes in the input patterns. After \nfurther training, the network learned to modify its predictions based on the input event \n(purely feed forward information) and, later still, on the context in which the event \noccurred (see Figure 2). Note that an important aspect of this type of recurrent network \nis that the representation of context is created by the network itself (Elman, 1989) and is \nnot specified a priori by the network designer or the environment (Jordan, 1986). In this \nway, the context can encode the information in the envirorunent which is relevant to the \nprediction task. Consequently, the network could, in principle, be adapted to other styles \nof music without modification of the design, or input and output representations. \n\na. Output vector \n__ ~--'\"'-_---'\"~_~~_----' L - _JLl ~ 1 - ' _~ __ \n\nn, \n\nb. Target histogram \n\nA A# B C C# D D# E F F# G G# 4 5 6 1 2 3 4 6 8 S W V A \n\ntones \n\noctaves \n\ndurations \n\naccent \n\nrest \n\nFigure 1: Comparison of output vector for the first event at Epoch 4 (a) and a histogram \nof the target vector averaged over all events (b). The upper grapb is the predicted output \nof Event 1 at Epoch 4. The lower grapb is a histogram of all the events in the piece, \ncreated by averaging all the target vectors. The comparison shows that the net learned \ninitially to predict the mean target before learning the variations specific to each event. \n\n\f1088 \n\nStevens and Wiles \n\n---.J1 ~ \n\n_______ --J,...'-___ --.J1 -11 ___ _ \n_________________ ~\n___________ ----JI\"L- -r-I1 --11 ..... , ___ _ \n_ ___ ___ ___ _ '\"'---\n---I'L-.n - - - I1 - - . . . . , . \n_____ __ ___ ---'~~._J1 ______ _ \n________ ~n~ ____ __ --...lL-.-Jl ______ _ \n__ ~n ..... __________ _ \n--..11'--__ _ \n\n--11..-\n\n11..- _ \n--11-\nrL- _ \n-.11___ _ \ni l - _ \n-11 __ \n-1L- _ \nI\"L..-\n_ \n\n=E~~~h~.~ _____ ~ _____ ~ ~ ---fl~_~_ \n-...J1'-___ _ \n---------------'\"'--- -.Jl~ \n---------------~ \n-------------------~ \n---------------~~ --I\"\\-.J\"1 ~ \n\n_---n --.fl'--___ _ \n\n___ .11 \n\n-..I1 ___ _ \n-..--J\"L..r, \n~ --..11 ___ _ \n~ ---11 ____ _ \n\n-----------~~ \n-------------~~ \n\n=E~~=h~2~ ______ ~ ____ ~'\"'___ ---l1 ----11'--__ _ \nM-.-.!'L-\n.. \n_ __ ___ ___ ___ __ ~n__ ---.J1 -11'--__ _ \n-1L- .. \n~ .. \n----~-----~-----''\"'--- ---l1 ~ \n_______ ~ ____ ~n__ ---J1 -1l'--__ _ \n-...1'l-\n______________ ~n__ --..J'l --1'1'--__ _ \n~ .. \n__ J'~_n -..I1'--__ _ \n~ .. \n- -n --11'---___ _ \n-..I1- .. \n.. ~ --\"''-----\n--J'L..-\n.. \nT..:..;~=-= ____ ~n,--____ ---.J1 -11'--__ _ \n_______ --lnl.-____ ---.J1 -11 _____ _ \n__________ 11- ~ ---1L--\n__________ n.---. ---Il -Il ______ _ \n---1L- ---1L--\n_ _________ fL.-\n__________ fL.----I\"L--Il ___ _ \n_______ ~n'_ ___ _ \n--..11 ______ _ \n__ --In'_ _________ _ ~\"'L- -11 ______ _ \n\n---I\"L-\n\n.. \n\n8 \n7 \n6 \n\n5 \u2022 3 \n\n5 \u2022 3 \n\n2 \n1 \n\n8 \n7 \n6 \n\n2 \n1 \n\n8 \n7 \n6 \n\n2 \n1 \n\n8 \n7 \n6 \n\n5 \u2022 3 \n\n5 \u2022 3 \n\n2 \n1 \n\nA M B C CI 0 Of E F F. Gat. 5 6 1 2 3 \u2022 6 8 S W V R event \n\ntones \n\noctaves \n\ndurations \n\naccent rest \n\nFigure 2: Evolution of the flrst eight events (predicted). The first block (targets) shows \nthe correct sequence of events for the four components. In the second block (Epoch 2), \nthe net is beginning to predict activation of strong and weak accents. In the third block \n(Epoch 4), the transition from one octave to another is evident. By the fourth block \n(Epoch 64), all four components are substantially correct. The pattern of activation \nacross the output vector can be interpreted as a statistical description of each musical \nevent. In psychological terms, the pattern of activation reflects the harmonic or chordal \nexpectancies induced by each note (Bharucha, 1987) and characteristic of the tonal \nhierarchy (Krumhansl & Shepard, 1979). \n\n3.1 NETWORK PERFORMANCE \n\nThe performance of the network as it learned to master The Blue Danube was recorded at \nlog steps up to 4096 epochs. One recording comprised the output predicted by the \nnetwork as the correct event was recycled as input to the network (similar to a teacher(cid:173)\nforcing paradigm). The second recording was generated by feeding the best guess of the \n\n \n~\n1\n \n~\n \n\fTonal Music as a Componential Code \n\n1089 \n\noutput back as input. The accent - the lilt of the waltz - was incorporated very early in \nthe training. Despite numerous errors in the individual events, the sequences were clearly \nidentifiable as phrases from The Blue Danube and, for the most part, errors in the tone \ncomponent were consistent with the tonality of the piece. The errors are of \npsychological importance and the overall performance indicated that the network learnt \nthe typical features of The Blue Danube and the waltz genre. \n\n3.2 \n\nINTERACTIONS BETWEEN ACCENT & DURATION \n\nWestern tonal music is characterised by regularities in pitch and timing components. For \nexample, the occurrence of particular tones and durations in a single composition is \nstructured and regular given that only a limited number of the possible combinations \noccur. Therefore, one way to gauge performance of the network is to compare the \nregularities extracted and represented in the model with the statistical properties of \ncomponents in the training composition. The expected frequencies of accent-duration \npairs, such as a quarter-note coupled with a strong accent, were compared with the actual \nfrequency of occurrence in the composition: the accent and duration couplings with the \nhighest expected frequencies were strong quarter-note (35.7), strong half-note (10.2), weak \nquarter-note (59.0), and weak half-note (16.9). Scrutiny of the predicted outputs of the \nnetwork over the time course of learning showed that during the initial training epochs \nthere was a strong bias toward the event with the highest expected frequency - weak \nquarter-note. The output of this accent-duration combination by the network decreased \ngradually. Prediction of a strong quarter-note by the network reached a value close to the \nexpected frequency of 35.7 by Epoch 2 (33) and then decreased gradually and approximated \nthe actual composition frequency of 19 at Epoch 64. Similarly, by Epoch 64, output of \nthe most common accent and duration pairs was very close to the actual frequency of \noccurrence of those pairs in the composition. \n\n3. 3 ANALYSIS OF HIDDEN UNIT ~EPRESENT A TIONS \n\nAn analysis of hidden unit space most often reveals structures such as regions, hierarchies \nand intersecting regions (Wiles & Bloesch, 1992). In the present network, four sub(cid:173)\nspaces would be expected (tone, octave, duration, accent), with events lying at the \nintersection of these suo-spaces. A two-dimensional projection of hidden unit space \nproduced from a canonical discriminant analysis (CDA) of duration-accent pairs reveals \nthese divisions (see Figure 3). In essence, there is considerable structure in the way \nevents are represented into clusters of regions with events located at the intersection of \nthese regions. In Figure 3 the groups used in the CDA relate to the output values which \nare observable groups. An additional CDA using groups based on position in bar showed \nthat the hidden unit space is structured around inferred variables as well as observable \nones. \n\n\f1090 \n\nStevens and Wiles \n\n2ft \n\n2ft \n\n21 \n\n\u2022 \n\nI, \n\n.. \n\n~ 1 \nJ \nI \n\nI \n\nI \n\n... 2w \n\n.. \n\n.. \n\n.. \n\nI \n\nFigure 3: Two-dimensional projection of hidden unit space generated by canonical \ndiscriminant analysis of duration-accent pairs. Each note in the composition is depicted \nas a labelled point, and the flrst and third canonical components are represented along the \nabscissa and ordinate. respectively. The first canonical component divides strong from \nweak accents (denoted s and w) . In the strong (s) accent region of the third canonical \ncomponent, quarter- and half-notes are separated (denoted by 2 and 4, respectively), and \nthe remaining right area separates rests, weak quarter- and weak half-notes. The \nsuperimposed line shows the first five notes of the opening two bars of the composition \nas a trajectory through hidden unit space: there is movement along the flrst canonical \ncomponent depending on accent (s or w) and the second bar starts in the half-note region. \n\n4 \n\nDISCUSSION & CONCLUSIONS \n\nThe focus of this study has been the extraction of information from the envirorunent(cid:173)\nthe temporal stream of events representing The Blue Danube - and its incorporation into \nthe static parameters of the weights and biases in the network. Evidence for the stages at \nwhich information from the environment is incorporated into the network representation \nis seen in the predicted output vectors (described above and illustrated in Figure 2). \nDifferent musical styles contain different kinds of infonnation in the components. For \nexample. the accent and duration components of a waltz take complementary roles in \nregulating the rhythm. From the durations of events alone, the position of a note in a \n\n\fTonal Music as a Componential Code \n\n1091 \n\nbar could be predicted without error. However, if the performer or listener made a single \nerror of duration, a rhythm system based on durations alone could not recover. By \ncontrast, accent is not a completely reliable predictor of the bar structure, but it is \neffective for recovery from rhythmic errors. The interaction between these two timing \ncomponents provides an efficient error correction representation for the rhythmic aspect of \nthe system. Other musical styles are likely to have similar regulatory functions \nperformed by different components. For example, consider the use of ornaments, such as \ntrills and mordents, in Baroque harpsichord music which, in the absence of variations in \ndynamics, help to signify the beat and metric structure. Alternatively, consider the \ninteraction between pitch and timing components with the placement of harmonically(cid:173)\nimportant tones at accented positions in a bar (Jones, 1992). \n\nThe network described here has learned transitions and relations between and within pitch \nand timing musical components and not simply the components per se. The interaction \nbetween accent and duration components, for example, demonstrates the manifestation of \na componential code in Western tonal music. Patterns of activation across the output \nvector represented statistical regularities or probabilities characteristic of the composition. \nNotably, the representation created by the network is reminiscent of the tonal hierarchy \nwhich reflects the regularities of tonal music and has been shown to be responsible for a \nnumber of performance and memory effects observed in both musically trained and \nuntrained listeners (Krumhansl, 1990). The distribution of activity across the tone \noutput units can also be interpreted as chordal or harmonic expectancies akin to those \nobserved in human behaviour by Bharucha & Stoeckig (1986). The hidden unit \nactivations represent the rules or grammar of the musical environment; an interesting \nproperty of the simple recurrent network is that a familiar sequence can be generated by \nthe trained network from the hidden unit activations alone. Moreover, the intersecting \nregions in hidden unit space represent composite states and the musical sequence is \nrepresented by transitions between states. Finally, the course of learning in the network \nshows an increasing specificity of predicted events to the changing context: during the \nearly stages of training, the default output or bias of the network is towards the average \npattern of activation across the entire composition but, over time, predictions are refined \nand become attuned to the pattern of events in particular contexts. \n\nAcknowledgements \nThis research was supported by an Australian Research Council Postdoctoral Fellowship \ngranted to the first author and equipment funds to both authors from the Departments of \nPsychology and Computer Science, University of Queensland. The authors wish to thank \nMichael Mozer for providing the musical database which was adapted and used in the present \nsimulation. The modification to McClelland & Rumelhart's (1989) bp program was developed \nby Paul Bakker, Department of Computer Science, University of Queensland. The comments \nand suggestions made by members of the Computer Science/Psychology Connectionist \nResearch Group at the University of Queensland are acknowledged. \n\nReferences \nBharucha. J. J. (1987). Music cognition and perceptual facilitation: A connectionist \n\nframework. Music Perception, 5, 1-30. \n\nBharucha, J. J. (1992). Tonality and learnability. \n\nIn M. R. Jones & S. Holleran (Eds.), \nCognitive bases of musical communication, pp. 213-223. WaShington: American \nPsychological Association . \n\n\f1092 \n\nStevens and Wiles \n\nBharucha, J. J., & Stoeckig, K. (1986). Reaction time and musical expectancy: Priming of \nchords. Journal of Experimental Psychology: Human Perception & Peiformance, 12, 403-\n410. \n\nBharucha, J., & Todd, P. M. (1989). Modeling the perception of tonal structure with neural \n\nnets. Computer Music Journal, 13, 44-53. \n\nBrousse, 0., & Smolensky, P. (1989). Virtual memories and massive generalization in \nIn Proceedings of the 11th Annual Conference of \n\nconnectionist combinatorial learning. \nthe Cognitive Science Society, pp. 380-387. Hillsdale, NJ: Lawrence Erlbaum. \n\nElman, J . L. (1989). Structured representations and connectionist models. (CRL Tech. Rep. \n\nNo. 8901). San Diego: University of California, Center for Research in Language. \n\nFodor, J. A., & Pylyshyn, Z. W. (1988). Connectionism and cognitive architecture: A critical \n\nanalysis. Cognition, 28, 3-71. \n\nHild, H., Feulner, J ., & Menzel, W . (1992). HARMONET: A neural net for harmonizing \nchorales in the style of J. S. Bach. In J. E. Moody, S. J. Hanson & R. Lippmann (Eds.), \nAdvances in Neural Information Processing Systems 4, pp. 267-274. San Mateo, CA.: \nMorgan Kaufmann. \n\nJones, M. R. (1992). Attending to musical events. \n\nIn M. R. Jones & S. Holleran (Eds.), \nCognitive bases of musical communication, pp. 91-110. Washington: American \nPsychological Association. \n\nJordan, M. I. (1986). Serial order: A parallel distributed processing approach (Tech. Rep. No. \n\n8604). San Diego: University of California, Institute for Cognitive Science. \n\nKrumhansl, C., & Shepard, R. N. (1979). Quantification of the hierarchy of tonal functions \nwithin a diatonic context. Journal of Experimental Psychology: Human Perception & \nPerformance, 5, 579-594. \n\nMcClelland, 1. L., & Rumelhart, D. E. (1989). Explorations in parallel distributed processing: \n\nA handbook of models, programs and exercises. Cambridge, Mass.: MIT Press. \n\nMonahan, C. B., Kendall, R. A ., & Carterette, E. C. (1987). The effect of melodic and \ntemporal contour on recognition memory for pitch change. Perception & Psychophysics, \n41, 576-600. \n\nMozer, M. C. (1991). Connectionist music composition based on melodic, stylistic, and \npsychophysical constraints. In P. M. Todd & D. G. Loy (Eds.), Music and connectionism, \npp. 195-211. Cambridge, Mass.: MIT Press. \n\nMozer, M. C. (1992). Induction of multiscale temporal structure. \n\nIn J. E. Moody, S. J. \nHanson, & R. P. Lippmann (Eds.), Advances in Neural Information Processing Systems 4, \npp. 275-282. San Mateo, CA.: Morgan Kaufmann. \n\nPhillips, S., & Wiles, J. (1993). Exponential generalizations from a polynomial number of \n\nexamples in a combinatorial domain . Submitted to IlCNN, Japan, 1993. \n\nPlaut, D. c., & McClelland, J. L. (1993). Generalization with componential attractors: Word \nand nonword reading in an attractor network. To appear in Proceedings of the 15th Annual \nConference of the Cognitive Science Society. Hillsdale, NJ: Erlbaum. \n\nRosenthal, D. (1992). Emulation of human rhythm perception. Computer Music Journal, 16, \n\n64-76. \n\nSayegh, S. (1989). Fingering for string instruments with the optimum path paradigm. \n\nComputer Music Journal, 13, 76-84. \n\nStevens, c., & Latimer, C. (1991). Judgments of complexity and pleasingness in music: The \n\neffect of structure, repetition, and training. Australian Journal of Psychology, 43, 17-22. \n\nStevens, C ., & Latimer, C . (1992). A comparison of connectionist models of music \n\nrecognition and human performance. Minds and Machines, 2, 379-400. \n\nTodd, P. M. (1988). A sequential network design for musical applications. In D. Touretzky, \nG . Hinton & T. Sejnowski (Eds.), Proceedings of the 1988 Connectionist Models Summer \nSchool, pp. 76-84. Menlo Park, CA: Morgan Kaufmann. \n\nWiles, J., & Bloesch, A. (1992). Operators and curried functions: Training and analysis of \nsimple recurrent networks. \nIn J. E. Moody, S. J. Hanson, & R . P. Lippmann (Eds.), \nAdvances in Neural Information Processing Systems 4. San Mateo, CA: Morgan Kaufmann. \n\n\f", "award": [], "sourceid": 759, "authors": [{"given_name": "Catherine", "family_name": "Stevens", "institution": null}, {"given_name": "Janet", "family_name": "Wiles", "institution": null}]}