{"title": "How Perception Guides Production in Birdsong Learning", "book": "Advances in Neural Information Processing Systems", "page_first": 110, "page_last": 116, "abstract": null, "full_text": "How Perception Guides  Production \n\n\u2022 In \n\nBirdsong Learning \n\nChristopher L.  Fry \ncfry@cogsci.ucsd.edu \n\nLa Jolla,  CA  92093-0515 \n\nDepartment of Cognitive Science \n\nUniversity  of California at San  Diego \n\nAbstract \n\nA  c.:omputational  model  of  song  learning  in  the  song  sparrow \n(M elospiza  melodia)  learns  to  categorize  the  different  syllables  of \na  song  sparrow  song  and  uses  this  categorization to  train  itself to \nreproduce  song.  The model fills  a crucial gap in the computational \nexplanation  of birdsong  learning  by  exploring  the  organization of \nperception  in  songbirds.  It shows  how  competitive  learning  may \nlead  to  the  organization  of  a  specific  nucleus  in  the  bird  brain, \nreplicates  the  song  production  results  of a  previous  model  (Doya \nand  Sejnowski,  1995),  and  demonstrates  how  perceptual  learning \ncan  guide production through  reinforcement  learning. \n\n1 \n\nINTRODUCTION \n\nThe  passeriformes  or  songbirds  make  up  more  than  half of all  bird  species  and \nare  divided  into  two  groups:  the  os cines  which  learn  their  songs  and  sub-oscines \nwhich do  not.  Oscines raised in isolation sing degraded species  typical songs similar \nto  wild  song.  Deafened  oscines sing  completely  degraded  songs  (Konishi,  1965)  , \nwhile  deafened  sub-oscines  develop  normal  songs  (Kroodsma  and  Konishi,  1991) \nindicating that auditory feedback  is  crucial in  oscine song learning. \n\nInnate structures in the bird  brain regulate song learning.  For example, song  spar(cid:173)\nrows show  innate preferences  for  their own species'  songs and song structure  (Mar(cid:173)\nler,  1991).  Innate preferences  are  thought to be encoded in an auditory template \nwhich  limits the  sounds  young  birds  may  copy.  According  to  the  auditory  tem(cid:173)\nplate  hypothesis  birds  go  through  two  phases  during  song  learning,  a  memo(cid:173)\nrization  phase and  a  motor phase.  In the  memorization phase, which  lasts \nfrom  approximately  20  to  50  days  after  birth  in  the  song sparrow,  the  bird selects \nwhich  sounds to copy  based  on  an innate  template  and refines  the  template  based \n\n\fHow Perception Guides Production in Birdsong Learning \n\n111 \n\nPOSTERIOR \n\nANTERIOR \n\n10 \u2022\u2022 chea \u2022  SJrIn I \n\n--+ ...... ing \n\nPall..,. \n\n~l'IodUclcn \n\nPall ... , \n\nFigure  1:  A simplified  sketch  of a  saggital  section of the songbird  brain.  Field  L  (Field \nL)  receives  auditory  input  and  projects  to  the  production  pathway:  HVc  (formerly  the \ncaudal  nucleus  of the hyperstriatum),  RA (robust  nucleus of archistriatum),  nXIIts (hy(cid:173)\npoglossal  nerve),  the syrinx  (vocal  organ)  and  the  learning  pathway:  X  (area  X),  DLM \n(medial  nucleus  of the  dorsolateral  thalamus),  LMAN  (lateral  magnocellular  nucleus  of \nthe  anterior neostriatum),  RA  (Konishi,  1989;  Vicario,  1994).  V  is  the lateral ventricle. \n\non  the sounds  it hears .  In  the motor phase (from  approximately 272  to 334  days \nafter  birth)  the  template  provides  feedback  during  singing.  Learning  to  sing  the \nmemorized,  template  song  is  a  gradual  process  of refining  the  produced  song  to \nmatch memory (Marler,  1991). \n\nA song is  made up of phrases, phrases  of syllables  and syllables of notes.  Syllables, \nusually  separated  by  periods  of silence,  are  the  main  units of analysis.  Notes  typ(cid:173)\nically  last  from  10-100  msecs  and  are  used  to  construct  syllables  (100-200  msecs) \nwhich  are  reused  to produce  trills  and other  phrases. \n\n2  NEUROBIOLOGY  OF  SONG \n\nThe two  main neural  pathways  that govern  song are  the motor  and learning path(cid:173)\nways  seen  in  figure  1  (Konishi ,  1989).  Lesions  to  the  motor  pathway  interrupt \nsinging  throughout  life  while  lesions  to  the  learning  pathway  disrupt  early  song \nlearning.  Although  these  pathways  seem  to  have  segregated  functions ,  recordings \nof neurons  during song playback have shown that cells  throughout the song system \nrespond  to song  (Konishi,  1989). \nStudies of song  perception have shown the best  auditory stimulus that will  evoke  a \nresponse  in  the  song  system  is  the  bird's  own  song  (Margoliash ,  1986) .  The song \nspecific  neurons  in HV c  of the  white-crowned  sparrow  often  require  a  sequence  of \ntwo syllables to respond  (Margoliash , 1986; Margoliash and  Fortune ,  1992)  and  are \nmade up of two main types in HV c .  One type is sensitive to temporal combinations \nof stimuli while  the  other  is  sensitive  to  harmonic characteristics  (Margoliash  and \nFortune,  1992) . \n\n3  COMPUTATION \n\nPrevious  computational work  on  birdsong learning  predicted  individual  neural  re(cid:173)\nsponses  using back-propagation (Margoliash and Bankes, 1993) and modelled motor \nmappings for  song  production  (Doya  and  Sejnowski,  1995).  The  current  work  de-\n\n\f112 \n\nC.L.FRY \n\n1  2 \n00 \n\n8 \n\nKohonen Neuron \n\nInpulLayer \n\n.. \n\nSliding  ~ _____  _ \nWindo\"\",. ~ \n\n1000 \n\nFigure 2:  Perceptual  network input  encoding.  The song  is  converted into frequency  bins \nwhich  are  presented  to  the Kohonen  layer over  four  time steps. \n\nvelops a  model of birdsong syllable perception which extends  Doya and Sejnowski's \n(1995)  model of birdsong learning.  Birdsong syllable segmentation is  accomplished \nusing  an  unsupervised  system  and  this system  is  used  to  train  the  network  to  re(cid:173)\nproduce  its  input using  reinforcement  learning. \n\nThe model implements the two phases  of the auditory template hypothesis,  mem(cid:173)\norization  and  motor.  In  the  first  phase  the  template  song  is  segmented  into \nsyllables  by  an  unsupervised  Kohonen  network  (Kohonen,  1984).  In  the  second \nphase  the  syllables  are  reproduced  using  a  reinforcement  learning paradigm based \non  Doya and Sejnowski  (1995). \n\nThe model extends  previous work in three  ways:  1)  a self-organizing network  picks \nout  syllables  in  the  song;  2)  the  self-organizing  network  provides  feedback  during \nsong production;  and 3)  a  more biologically plausible model of the syrinx is  used  to \ngenerate  song. \n\n3.1  Perception \n\nRecognizing  a  syllable  involves  identifying  a  short  sequence  of  notes.  Kohonen \nnetworks  use  an  unsupervised  learning method  to categorize  an  input space  based \non  similar  neural  responses.  Thus  a  Kohonen  network  is  a  natural  candidate  for \nidentifying the syllables in  a  song. \n\nOne  song  from  the  repertoire  of a  song  sparrow  was  chosen  as  the  training  song \nfor  the  network.  The  song  was  encoded  by  passing  a  sliding  window  across  the \ntraining waveform (sampled at 22 .255 kHz)  of the selected  song.  At each time step, \na  non-overlapping 256  point  (~ .011  sec)  fast  fourier  transform (FFT)  was  used  to \ngenerate  a power spectrum  (figure  2).  The power spectrum was  divided into 8 bins. \nEach bin was  mapped to a real number using a gaussian summation procedure with \nthe  peak of the gaussian at  the center  of each  frequency  bin.  Four  time-steps  were \npassed  to each  Kohonen  neuron. \n\nThe  network's  task  was  to  identify  similar syllables  in  the  input  song.  The  input \nsong  was  broken  down  into  syllables  by  looking for  points  where  the  power  at  all \n\n\fHow Perception Guides  Production in Birdsong Learning \n\n113 \n\n10 \n\n>. u \n.: \n\" g.  5 \n\" \" ... \n\nt \n\n-/>, \n\na \nkH<  s \n\n0.0 \n\nn1 \n\nCD ..  n3 \n\"  n2 \nii: \n..  n5 \nS  n4 \n:I  n6 \nCD z  n7 \nn8 \n\n0.5 \n\n1.0 \ntime \n\n20 \n\nFigure 3:  Categorization of song  syllables by  a Kohonen  network.  The power-spectrum of \nthe  training song  is  at  the top.  The  responses  of the  Kohonen  neurons  are  at  the  bottom. \nFor  each  time-step  the  winning  neuron  is  shown  with  a  vertical  bar.  The  shaded  areas \nindicate  the  neuron  that fired  the most  during  the  presentation  of the syllable. \n\nfrequencies  dropped below  a threshold .  A syllable was  defined  as  sound of duration \ngreater  than .011  seconds  bounded by two  low-power  points .  The network  was  not \ntrained  on  the  noise  between  syllables.  The  song  was  played  for  the  network  ten \ntimes  (1050  training vectors),  long enough for  a stable response  pattern to emerge. \nThe  activation of a  neuron  was:  N etj  = 'ExiWij'  Where:  N etj  = output of neuron \nj , Wij  =  the  weight  connecting inputi  to  n euronj ,  Xi  =  inputi.  The  Kohonen  net(cid:173)\nwork  was  trained by initializing the connection weights  to  1/Jnumber  of neurons \n+ small  random  component  (r  S;  .01) ,  normalizing  the  inputs ,  and  updating  the \nweights  to  the  winning neuron  by  the  following  rule :  W n ew  =  W old  + a(x - W old) \nwhere :  a  =  training rat e  =  .20 .  If the same neuron won twice in a row the train(cid:173)\ning rate  was  decreased  by  1/2.  Only  the  winning  neuron  was  reinforced  resulting \nin  a  non-localized feature  map . \n\n3.1.1  Perceptual Results \n\nThe  Kohonen  network  was  able  to assign  a  unique  neuron  to  each  type  of syllable \n(figure 3) .  Of the eight neurons in the network. the one that fired the most frequently \nduring  the  presentation  of a  syllable  uniquely  identified  the  type  of syllable.  The \nfirst  four  syllables  of the  input  song  sound  alike,  contain  similar frequencies ,  and \nare  coded  by  the  first  neuron  (N1).  The  last  three  syllables  sound  alike,  contain \nsimilar  frequencies ,  and  are  coded  by  the  fourth  neuron  (N4).  Syllable  five  was \ncoded  by  neuron  six  (N6) , syllable  six  by  neuron  two  (N2)  and  syllable  seven  by \nneuron  eight  (N8). \nFigure 4 shows the frequency sensitivity of each neuron (1-8, figure 3) plotted against \neach  time step  (1-4).  This  plot shows  the  harmonic  and  temporally sensitive  neu(cid:173)\nrons  that  developed  during the  learning phase  of the  Kohonen  network.  Neuron  2 \nis  sensitive  to  only one frequency  at  approximately 6-7  kHz , indicated  by  the solid \nwhite  band  across  the  6-7  kHz  frequency  range  in  figure  4.  Neuron  4  is  sensitive \nto  mid-range  frequencies  of short  duration.  Note  that  in  figure  4  N4  responds \n\n\f114 \n\nC. L. FRY \n\nN5 \n\no  1 \n\n2\nN6 \n\n3\n\n4 \n\n01 23 4 \n\nN7 \n\nN8 \n\nTime  S t e p \n\nFigure  4:  The  values  of the  weights  mapping  frequency  bins  and  time steps  to  Kohonen \nneurons.  White is  maximum, Black is  minimum. \n\nmaximally  to  mid-range  frequencies  only  in  the  first  two  time  steps.  It uses  this \ntemporal sensitivity to distinguish between the last three syllables and the fifth  syl(cid:173)\nlable (figure  3)  by keying off the  length of time mid-range frequencies  are  present. \nContrast  this  early  response  sensitivity  with  neuron  6,  which  is  sensitive  to  mid(cid:173)\nrange  frequencies  of long  duration ,  but  responds  only  after one  time step .  It uses \nthis temporal sensitivity to respond to the long sustained frequency  of syllable four . \nConsidered  together,  neurons  2,4,6 and 8 illustrate the two  types  of neurons  (tem(cid:173)\nporal and harmonic) found in HVc by Margoliash and Fortune (1993).  Competitive \nlearning may  underly  the formation of these  neurons  in  HV c. \n\n3.2  Production \n\nAfter  competitive learning trains  the  perceptual  part  of the  network  to  categorize \nthe song into syllables, the  perceptual network can be  used  to train the production \nside  of the network  to sing. \n\nThe first  step  in  modelling  song  production  is  to  create  a  model  of the  avian  vo(cid:173)\ncal  apparatus ,  the  syrinx.  In  the  syrinx sounds  arise  when  air  flows  through  the \nsyringeal  passage  and  causes  the  tympanic membrane  to vibrate.  The frequency  is \ncontrolled by  the tension of the membrane controlled by the syringeal musculature. \nThe amplitude is  dependent  on the area of the syringeal orifice  which  is  dependent \non  the  tension  of the  labium.  The  interactions  of this  system  were  modelled  by \nmodulated  sine  waves.  Four  parameters  governed  the  fundamental  frequency(p) , \nfrequency  modulation(tm) ,  amplitude  (ex)  and  frequency  of amplitude  modula(cid:173)\ntion(I).  The range of the parameters was set  according to calculations in  Greenwalt \n(1968).  The parameters  were  combined in the following  equation  (based  on Green(cid:173)\nwalt,  1968),  f(ex , l,p, tm , t) = excos(21l\"t  1) cos(21l\"t  p + cos(21l\"t  tm)) . \nUsing this equation song can be generated over time by making assumptions about \nthe  response  properties  of neurons  in  RA . Following Doya and Sejnowski  (1995)  it \nwas  assumed  that  pools  of RA  neurons  have  different  temporal response  profiles. \nSyllable  like  temporal  responses  can  be  generated  by  modifying  the  weights  from \nthe  Kohonell  layer  (HV c)  to the production layer  (RA) . \n\n\fHow Perception Guides Production in Birdsong Learning \n\n115 \n\nTnnni~ Song \n\n,  .f. \u00b7~:if>vr. \n... \n\n. ..  . \n\ni \n'0 \nTim. \n\ni \n\n15 \n\nI \n\n20 \n\nJ'etworl< Song trained with Spectmgmm Target \n\nTi'llll! \n\nJ'etworl< Song trained with J'euraJ.  Activation Target \n\n10 \n\nkHz \n\nS \n\na a \n\nas \n\n10 \n\n15 \n\n20 \n\nFigure  5:  Training  song  and  two  songs  produced  with  different  representations  of  the \ntraining song. \n\nThe  production  side  of the  network  was  trained  using  the  reinforcement  learning \nparadigm described  in  Doya and Sejnowski  (1995).  Each  syllable was  presented  in \nthe order  it occurred  in the training song  to the  Kohonen  layer,  which  turned  on a \nsingle neuron.  A random vector was added to the weights from the Kohonen layer to \nthe output layer  and a syllable was produced.  The produced syllable was  compared \nto  the  stored  representation  of the  template  song  which  was  used  to  generate  an \nerror  signal  and  an  estimate  of the  gradient.  If  the  evaluation  of the  produced \nsyllable  was  better  than  a  threshold  the  weights  were  kept,  otherwise  they  were \ndiscarded. \n\nTwo  experiments  were  done  using  different  representations  of the  template  song. \nIn  the  first  experiment  the  template song  was  the  stored  power  spectrum  of each \nsyllable and the error signal was the cosine of the angle between the power spectrum \nof the  produced  syllable  and  the  template  syllable.  In  the second  experiment  the \ntemplate song was  the stored  neural  responses  to song  (recorded  during the mem(cid:173)\norization  phase)  and  the  error  signal  was  the  Euclidean  distance  between  neural \nresponses  to the produced  syllable and  the neural  responses  to the template song. \n\n3.2.1  Production Results \n\nFigure  5  shows  the output  of the  production  network  after  training  with  different \nrepresentations  of the  training song.  The  network  was  able  to replicate  the  major \nfrequency  components of the training song  to a  high  degree  of accuracy.  The song \ntrained  with  the spectrogram target  was  learned  to  a  90%  average  cosine  between \nthe spectrograms of the produced song  and the training song on each syllable with \nthe best syllable learned to 100% accuracy and the worst to 85% after 1000 trials.  A \ncrucial  aspect  to achieving performance  was  smoothing the template spectrogram. \nThe third song shows that the network was able to learn the template song using the \nneural responses of the perceptual system to generate the reinforcement signal.  The \naverage  distance  between  the initial randomly produced  syllables  and  the  training \n\n\f116 \n\nsong was  reduced  by  50%. \n\n4  DISCUSSION \n\nC. L.FRY \n\nThis work fills  a  crucial  gap in  the  computational explanation of song  learning left \nby  prior  work .  Doya  and  Sejnowski  (1995)  showed  how  song  could  be  produced \nbut left  unanswered  the questions  of how song is  perceived  and how  the perceptual \nsystem  provides  feedback  during  song  production.  This  study  shows  a  time-delay \nKohonen  network  can  learn  to  categorize  the  syllables  of a  sample  song  and  this \nnetwork can train song production with no external teacher.  The Kohonen  network \nexplains  how  neurons  sensitive  to  temporal  and  harmonic structure  could  arise  in \nthe  songbird  brain  through  competitive  learning.  Taken  as  a  whole ,  the  model \npresents  a  concrete  proposal of the  computational principles  governing the  Audi(cid:173)\ntory Template Hypothesis and  how  a song is  memorized  and used  to train song \nproduction.  Future work will flesh  out the effects  of innate structure on learning by \nexamining how the settings of the initial weights on the network affect song learning \nand  predict experimental effects  of deafening  and  isolation. \n\nAcknowledgements \n\nThanks to S.  Vehrencamp for  providing the song data, J . Batali, J. Elman, J.  Brad(cid:173)\nbury and T. Sejnowski for  helpful comments, and K.  Doya for  advice on replicating \nhis  model. \n\nReferences \n\nDoya,  K . and Sejnowski, T .J.  (1995).  A novel reinforcement model of bird song vocalization \nlearning.  In  Tesauro,  G .,  Touretzky,  D.  S.  and  Leen ,  T.K.,  editors,  Advances  in  Neural \nInformation  Processing Systems  7.  MIT Press,  Cambridge,  MA. \n\nGreenwalt,  C.H.  (1968).  Bird  Song:  Acoustics  and  Physiology.  Smithsonian  Institution \nPress.  Wash.,  D.C. \n\nKohonen,  T .  (1984).  Self-organization  and Associative  Memory,  Vol.  8.  Springer-Verlag, \nBerlin. \n\nKonishi,  M.  (1965).  The  role  of  auditory  feedback  in  the  control  of  vocalization  in  the \nwhite-crowned  sparrow.  Zeitschrijt fur  Tierpsychogie , 22,770-783. \n\nKonishi,  M.  (1989).  Birdsong for  Neurobiologists.  Neuron, 3,  541-549. \n\nKroodsma,  D.E .  and  Konishi ,  M.  (1991) .  A  suboscine  bird  (eastern  phoebe,  Sayonoris \nphoebe)  develops  normal  song  without  auditory feedback.  Animal Behavior, 42, 477-487. \n\nMarler,  P.  (1991).  The instinct  to  learn.  In  The  Epigenesis  of Mind:  Essays  on  Biology \nand  Cognition,  eds.  S.  Carey  and  R.  Gelman.  Lawrence  Erlbaum  Associates. \n\nMargoliash ,  D .  (1986).  Preference  for  autogenous  song  by  auditory  neurons  in  a  song \nsystem  nucleus  of the  white-crowned sparrow .  Journal of Neuroscience, 6,1643-1661. \n\nMargoliash,  D .  and  Bankes,  S.C.  (1993) .  Computations in  the Ascending  Auditory  Path(cid:173)\nway  in  Songbirds  Related  to  Song  Learning.  American Zoologist, 33,  94-103. \n\nMargoliash ,  D.  and  Fortune,  E.  (1992).  Temporal  and  Harmonic  Combination-Sensitive \nNeurons  in the  Zebra  Finch's  HVc.  Journal of Neuroscien ce,  12,  4309-4326. \n\nVicario ,  D.  (1994).  Motor  Mechanisms  Relevant  to  Auditory-Vocal  Interactions in  Song(cid:173)\nbirds.  Brain,  Behavior and Evolution,44,  265-278 . \n\n\f", "award": [], "sourceid": 1069, "authors": [{"given_name": "Christopher", "family_name": "Fry", "institution": null}]}