{"title": "Causal Categorization with Bayes Nets", "book": "Advances in Neural Information Processing Systems", "page_first": 99, "page_last": 105, "abstract": null, "full_text": "Causal  Categorization  with  Bayes  Nets \n\nBob Rehder \n\nDepartment of Psychology \n\nNew  York University \nNew  York, NY  10012 \nbob .rehder@nyu.edu \n\nAbstract \n\nA  theory  of  categorization  is  presented  in  which  knowledge  of \ncausal  relationships  between  category  features  is  represented  as  a \nBayesian  network.  Referred  to  as  causal-model  theory,  this  theory \npredicts  that  objects  are  classified  as  category  members  to  the \nextent they  are  likely  to  have  been  produced  by  a categorys  causal \nmodel.  On  this  view,  people  have  models  of  the  world  that  lead \nthem  to  expect  a  certain  distribution  of  features \nin  category \nmembers  (e.g.,  correlations  between  feature  pairs  that  are  directly \nconnected  by  causal  relationships),  and  consider  exemplars  good \ncategory  members  when  they  manifest  those  expectations.  These \nexpectations  include  sensitivity  to  higher-order  feature  interactions \nthat emerge from the  asymmetries inherent in causal relationships. \n\nResearch  on  the  topic  of categorization  has  traditionally  focused  on  the  problem  of \nlearning  new  categories  given  observations  of  category  members.  In  contrast,  the \ntheory-based  view  of  categories  emphasizes  the  influence  of  the  prior  theoretical \nknowledge  that  learners  often  contribute  to  their  representations  of  categories  [1]. \nHowever, in  contrast to  models  accounting  for  the  effects  of empirical  observations, \nthere  have  been  few  models  developed  to  account for  the  effects of prior knowledge. \nThe  purpose  of  this  article  is  to  present  a  model  of  categorization  referred  to  as \ncausal-model  theory  or  CMT  [2,  3].  According  to  CMT,  people 's  know ledge  of \nmany  categories  includes  not only  features,  but  also  an  explicit representation  of the \ncausal mechanisms that people believe link the features  of many categories. \n\nIn  this  article  I  apply  CMT  to  the  problem  of  establishing  objects  category \nmembership.  In  the  psychological  literature  one  standard  view  of  categorization  is \nthat  objects  are  placed  in  a  category  to  the  extent they  have  features  that  have  often \nbeen  observed  in  members  of that category.  For example, an  object that has  most of \nthe  features  of  birds  (e.g.,  wings,  fly,  build  nests  in  trees,  etc.)  and  few  features  of \nother  categories  is  thought to  be  a  bird. This  view  of categorization  is  formalized  by \nprototype  models  in  which  classification  is  a  function  of the  similarity  (i.e. , number \nof  shared  features)  between  a  mental  representation  of  a  category  prototype  and  a \nto-be-classified  object.  However ,  a  well-known  difficulty  with  prototype  models  is \nthat  a  features  contribution  to  category  membership  is  independent  of the  presence \nor  absence  of  other  features.  In  contrast ,  consideration  of  a  categorys \ntheoretical \ninfluence  which  combinations  of  features  make  for \nknowledge \nacceptable  category  members.  For  example ,  people  believe  that  birds  have  nests  in \ntrees  because  they  can  fly , and  in  light  of this  knowledge  an  animal  that  doesnt  fly \n\nlikely \n\nis \n\nto \n\n\fand  yet  still  builds  nests  in  trees  might  be  considered  a  less  plausible  bird  than  an \nanimal  that  builds  nests  on  the  ground  and  doesnt  fly  (e.g.,  an  ostrich)  even  though \nthe latter animal has fewer features  typical of birds. \n\nTo  assess  whether  knowledge  in  fact  influences  which  feature  combinations  make \nfor  good  category  members , in  the  following  experiment undergraduates  were  taught \nnovel  categories  whose  four  binary  features  exhibited  either  a  common-cause  or  a \ncommon-effect  schema  (Figure  1).  In  the  common-cause  schema,  one  category \nfeature  (PI)  is  described  as  causing  the  three  other  features  (F2,  F3,  and  F4).  In  the \ncommon-effect  schema  one  feature  (F4)  is  described  as  being  caused  by  the  three \nothers  (F I,  F2,  and  F3).  CMT  assumes  that  people  represent  causal  knowledge  such \nas  that  in  Figure  1  as  a  kind  of Bayesian  network  [4]  in  which  nodes  are  variables \nrepresenting  binary  category  features  and  directed  edges  are  causal  relationships \nrepresenting  the  presence  of  probabilistic  causal  mechanisms  between  features. \nSpecifically ,  CMT  assumes  that  when  a  cause  feature  is  present  it  enables  the \noperation  of a  causal  mechanism  that  will,  with  some  probability  m , bring  about the \npresence  of the  effect feature.  CMT  also  allow  for  the  possibility  that effect features \nhave  potential background causes  that are  not explicitly  represented  in  the  network, \nas  represented  by  parameter  b  which  is  the  probability  that  an  effect  will  be  present \neven  when  its  network causes  are  absent.  Finally, each  cause  node  has  a parameter c \nthat represents the probability that a cause feature  will be  present. \n\n~ \n\n~ \u00ae \n\nCommon-Effect \n\n. (~~)  @ . \n.. \n.... \n\" .... @/  \u00ae'. \n\n~~:f\u00b7\"\"\"\"\u00ae1  \u00ae\"\"\"::\u00ae \n\nF \n3 \n\n.\u2022 \n\n: \n\nSchema \n\nCorrelations \n\nCommon-Cause \n\nCommon-Effect \n\nCorrelations \n\nCommon-Cause \n\nSchema \n\nFigure 1. \n\nFigure 2. \n\nThe  central  prediction  of  CMT  is  that  an  object  is  considered  to  be  a  category \nmember  to  the  extent  that  its  features  were  likely  to  have  been  generated  by  a \ncategory's  causal  mechanisms.  For  example,  Table  1  presents  the  likelihoods  that \nthe  causal  models  of Figure  1  will  generate  the  sixteen  possible  combinations  of F I, \nF2, F3, and  F4. Each  likelihood  equation  can  be  derived  by  the  application  of simple \nBoolean  algebra  operations.  For  example,  the  probability  of exemplar  1101  (F I, F2, \nF4  present, F3  absent)  being  generated  by  a  common-cause  model  is  the  probability \nthat  F I  is  present  [c],  times  the  probability  that  F2  was  brought  about  by  F I  or  its \nbackground  cause  [1- (lmj(l-b)],  times  the  probability  that  F3  was  brought  about \nby  neither  F I  nor  its  background  cause  [(l-m )(l-b)],  times  the  probability  that  F 4 \nwas  brought  about  by  F I  or  its  background  cause  [1- (lmj(l-b)].  Likewise ,  the \nprobability  of exemplar  1011 \n(F I,  F3,  F 4 present,  F2  absent)  being  generated  by  a \ncommon-effect  model  is  the  probability  that  FI  and  F3  are  present  [c 2 ],  times  the \nprobability  that  F2  is  absent  [1-\u00a3],  times  the  probability  that  F4  was  brought  about \nby  F I, F3, or its  background  cause  [1- (lmj(l-m )(l-b)] .  Note  that these  likelihoods \nassume  that  the  causal  mechanisms  in  each  model  operate  independently  and  with \nthe same probability m, restrictions that can be relaxed in  other applications. \n\nThis  formalization  of  categorization  offered  by  CMT \nthat  peoples \ntheoretical  knowledge  leads  them  to  expect  a  certain  distribution  of  features  in \ncategory  members ,  and  that  they  use  this  information  when  assigning  category \nmembership.  Thus , to  gain  insight  into  the  categorization  performance  predicted  by \nCMT ,  we  can  examine  the  statistical  properties  of  category  features  that  one  can \n\nimplies \n\n\fexpect  to  be  generated  by  a  causal  model.  For  example ,  dotted  lines  in  Figure  2 \nrepresent  the  features  correlations  that  are  generated  from  the  causal  schemas  of \nFigure  1.  As  one  would  expect,  pairs  of  features  directly \nlinked  by  causal \nthe  common-cause  schema  F I  is  correlated  with  its \nrelationships  are  correlated  in \neffects  and  in  the  common-effect  schema  F4  is  correlated  with  its  causes.  Thus, \nCMT  predicts \nthat  combinations  of  features  serve  as  evidence  for  category \nmembership  to  the  extent  that  they  preserve  these  expected  correlations  (i.e. ,  both \ncause  and  effect  present  or  both  absent) ,  and  against  category  membership  to  the \nextent that they  break those correlations (one present and  the other absent). \n\nTable  1:  Likelihoods Equations and Observed and Predicted Values \n\nCommon Cause Schema \n\nLikelihood \n\nObserved Predicted \n\nCommon Effect Schema \n\nControl \nObserved Predicted Observed \n\ne'b ,3 \ne 'b ,2 b \ne'b ,2 b \ne 'b ,2 b \ne m ,3 b ,3 \ne 'b 'b 2 \ne'b 'b 2 \ne 'b 'b 2 \n\nExemElar \n60.0 \n0000 \n44.9 \n0001 \n46.1 \n0010 \n42.8 \n0100 \n44.5 \n1000 \n41.0 \n0011 \n0101 \n40.8 \n42.7 \n0110 \n1001 \n55.1 \n52.6 \n1010 \n54.3 \n1100 \n0111 \n39.4 \n1011 \n64.2 \n1101 \n65 .3 \n1110 \n62.0 \n90.8 \n1111 \nNote . e'=l- c .  m '=l-m .  b'=l-b. \n\ne m ,2 b ,2 (1- m 'b ') \nem ,2 b ,2 (1- m 'b ') \nem ,2 b ,2 (1- m 'b ') \n\nem 'b '(1-m 'b ,)2 \ne m 'b '(1-m 'b ,)2 \nem 'b '(1-m 'b ,)2 \n\ne (1-m 'b ,)3 \n\ne 'b 3 \n\nLikelihood \n\ne ,3 b , \ne ,3 b \n\nee,2 m 'b ' \nee ,2 m 'b ' \nee,2 m 'b ' \n\nee ,2 (1-m 'b ') \nee,2 (1-m 'b ') \ne 2e 'm ,2 b , \nee,2 (1-m 'b ') \ne 2e 'm ,2 b , \ne 2e 'm ,2 b , \n\ne 2e'(1-m ,2 b ,) \ne 2e '(1-m ,2 b ,) \ne 2e'(1-m ,2 b ,) \n\ne 3m ,3 b , \n\ne 3(1-m ,3 b ,) \n\n61.7 \n45 .7 \n45 .7 \n45 .7 \n44.1 \n40.1 \n40.1 \n40.1 \n52.7 \n52 .7 \n52 .7 \n38.1 \n65 .6 \n65 .6 \n65 .6 \n89 .6 \n\n70.0 \n26.3 \n43.4 \n47 .3 \n48.0 \n56.3 \n56.5 \n38 .3 \n57.7 \n43 .0 \n41.9 \n71.0 \n75 .7 \n74.7 \n33 .8 \n91.0 \n\n69 .3 \n27 .8 \n47 .7 \n47 .7 \n47 .7 \n56.5 \n56.5 \n39 .2 \n56.5 \n39 .2 \n39 .2 \n74.4 \n74.4 \n74.4 \n35 .8 \n90 .0 \n\n70 .7 \n67.0 \n65.6 \n66.0 \n67.0 \n67.1 \n66.5 \n65.6 \n68.0 \n67.6 \n69 .9 \n67.6 \n67 .2 \n70 .2 \n72 .2 \n75.6 \n\nCausal  networks  not  only  predict  pairwise  correlations  between  directly  connected \nfeatures.  Figure  2  indicates  that  as  a  result  of  the  asymmetries  inherent  in  causal \nrelationships  there  is  an  important  disanalogy  between  the  common-cause  and \ncommon-effect  schemas:  Although  the  common-cause  schema  implies  that  the  three \neffects  (F2 ,  F3 ,  F4)  will  be  correlated  (albeit  more  weakly  than  directly  connected \nfeatures) ,  the  common-effect  schema  does  not  imply  that  the  three  causes  (F I ,  F2 , \nF3)  will  be  correlated.  This  asymmetry  between  common-cause  and  common-effect \nschemas  has  been  the  focus  of  considerable  investigation  in  the  philosophical  and \npsychological  literatures  [3 ,  5].  Use  of  these  schemas  in  the  following  experiment \nenables  a  test  of  whether  categorizers  are  sensitive  the  pattern  of  correlations \nbetween  features  directly-connected  by  causal  laws,  and  also  those  that  arise  due  to \nthe  asymmetries  inherent in  causal relationships  shown  in  Figure  2.  Moreover , I  will \nshow  that  CMT  predicts,  and  humans  exhibit,  sensitivity  to  interactions  among \nfeatures  of a higher-order than  the  pairwise interactions shown in  Figure 2. \n\nMethod \n\nSix  novel  categories  were  used  in  which  the  description  of  causal  relationships \nbetween  features  consisted  of  one  sentence  indicating  the  cause  and  effect  feature , \nand  then  one  or  two  sentences  describing  the  mechanism  responsible  for  the  causal \nrelationship.  For  example ,  one  of  the  novel  categories ,  Lake  Victoria  Shrimp ,  was \ndescribed  as  having  four  binary  features \n(e.g. ,  A  high  quantity  of  ACh \nneurotransmitter.  ,  Long-lasting  flight  response.  ,  Accelerated  sleep  cycle.  ,  etc.) \n\n\fand  causal  relationships  among  those  features  (e.g. ,  \"A  high  quantity  of  ACh \nneurotransmitter  causes  a  long-lasting  flight  response.  The  duration  of the  electrical \nsignal to  the muscles is  longer because of the excess amount of neurotransmitter. \"). \n\nfour \n\nin \n\nfeatures.  Participants \n\nParticipants  first  studied  several  computer  screens  of  information  about  their \nassigned  category  at  their  own  pace.  All  participants  were  first  presented  with  the \nthe  common-cause  condition  were \ncategorys \nadditionally  instructed  on  the  common-cause  causal  relationships  (F 1-;' F2 ,  F 1-;' F3 , \nF 1-;' F 4) ,  and  participants  in  the  common-effect  condition  were  instructed  on  the \ncommon-effect  relationships  (F 1-;.F4 ,  F2-;.F4 ,  F3-;.F4 ).  When  ready ,  participants \ntook  a  multiple-choice  test  that  tested  them  on  the  knowledge  they  had  just studied. \nParticipants  were required to  retake  the  test until they  committed 0  errors. \n\nParticipants  then  performed  a  classification  task  in  which  they  rated  on  a  0-100 \nscale  the  category  membership  of  16  exemplars ,  consisting  of  all  possible  objects \nthat  can  be  formed  from  four  binary  features.  For  example ,  those  participants \nassigned  to  learn  the  Lake  Victoria  Shrimp  category  were  asked  to  classify  a  shrimp \nthat  possessed  \"High  amounts  of  the  ACh  neurotransmitter ,\"  \"A  normal  flight \nresponse ,\"  \"Accelerated  sleep  cycle ,\"  and  \"Normal  body  weight.\"  The  order  of  the \ntest exemplars was randomized for each participant. \n\nOne  hundred  and  eight  University  of  Illinois  undergraduates  received  course  credit \nfor  participating  in  this  experiment.  They  were  randomly  assigned  in  equal  numbers \nto  the  three  conditions , and to  one of the  six experimental categories. \n\nResults \n\nCategorization  ratings  for  the  16  test  exemplars  averaged  over  partIclpants  in  the \ncommon-cause ,  common-effect,  and  control  conditions  are  presented  in  Table  1. \nThe  presence  of  causal  knowledge  had  a  large  effect  on  the  ratings.  For  instance, \nexemplars  0111  and  0001  were  given  lower  ratings  in  the  common-cause  and \ncommon-effect conditions , respectively  (39.4  and  26.3)  than  in  the  control condition \n(67.6  and  67.0)  presumably  because  in  these  exemplars  correlations  are  broken \n(effect  features  are  present  even  though  their  causes  are  absent).  In  contrast, \nexemplar  1111  received  a  significantly  higher  rating  in  the  common-cause  and \ncommon-effect  conditions  than  in  the  control  condition  (90.8  and  9l.0  vs.  75.6) , \npresumably  because in  both conditions all correlations are preserved. \n\nTo  confirm  that  causal  schemas  induced  a  sensitivity  to  interactions  between \nfeatures,  categorization  ratings  were  analyzed  by  performing  a  multiple  regression \nfor  each  participant.  Four  predictor  variables  (f1 , f2,  f3 , f4)  were  coded  as  -1  if the \nfeature  was  absent,  and  + 1  if it  was  present.  An  additional  six  predictor  variables \nwere  formed  from  the  multiplicative  interaction  between  pairs  of features:  f12 ,  f13 , \nf14 , f24 , f34 , and  f23.  For  those  feature  pairs  connected  by  a  causal relationship  the \ntwo-way  interaction  terms  represent  whether  the  causal  relationship  is  confirmed \n(+ 1,  cause  and  effect  both  present  or  both  absent)  or  violated  (-1 ,  one  present  and \none  absent).  Finally , the  four  three-way  interactions  (f123 , f124 , f134,  and  f234) , and \nthe  single four-way  interaction (f1234)  were  also  included  as  predictors. \n\nRegression  weights  averaged  over  participants  are  presented  in  Figure  3  as  a \nfunction  of  causal  schema  condition.  Figure  3  indicates  that  the  interaction  terms \ncorresponding  to  those  feature  pairs  assigned  causal  relationships  had  significantly \npositive  weights  in  both  the  common-cause  condition  (f12 ,  f13 ,  f14) ,  and  the \ncommon-effect condition  (f14 , f24 , f34).  That is , as  predicted  (Figure 2)  an  exemplar \nwas  rated  a  better  category  member  when  it  preserved  expected  correlations  (cause \nand  effect  feature  either  both  present  or  both  absent) , and  a  worse  member  when  it \nbroke those correlations (one absent  and the other present). \n\n\f(a)  Common Cause vs. Control \n\n\u2022 \n\nCC Observed \n\n12 \n\n10 \n\n6 \n\n8 \n\n.l: \nOf) \n'0:; \n~ \n~  4 \na \n'\" \n'\" \"'\" \n\n2 \n\n0 \n\nControl Observed \n\nCC Predicted \n\nE9 \n~ \n\n\u2022 \n\nCE Observed \n\nControl Observed \n\nCE Predicted \n\n(2) \n12 \n\n10 \n\n(b) Common Effect vs. Control \n\n8 \n\n.l: \nOf) \n'0:; \n6 \n~ \n~  4 \na \n'\" '\" \n\"'\" \n\n2 \n\n0 \n\n(2) \n\nfl \n\nf2 \n\nf3 \n\nf4 \n\nfl2 \n\nfl3 \n\nf23  fl23  f124  f134  f234  f1234 \n\nf24 \n\nfl4 \nf34 \nRegression Term \n\nFigure 3 \n\nthan  directly-linked  features.  Consistent  with \n\nIn  addition,  it  was  shown  earlier  (Figure  2)  that  because  of their  common-cause  the \nthree  effect  features  in  a  common-cause  schema  will  be  correlated,  albeit  more \nweakly \nthis \ncondition  the  three  two-way  interaction  terms  between  the  effect  features  (f24,  f34, \nf23)  are  greater  than  those  interactions  in  the  control  condition.  In  contrast,  the \ncommon-effect  schema  does  not  imply \nthree  cause  features  will  be \ncorrelated,  and  in  fact  in  that condition  the  interactions  between  the  cause  attributes \n(f12, f13, f23)  did  not differ from those in  the control condition (Figure  3). \n\nthis  prediction,  in \n\nthat  the \n\nFigure  3  also  reveals  higher-order  interactions  among  features  in  the  common-effect \ncondition:  Weights  on  interaction  terms  f124,  f134,  f234,  and  f1234  (- 1.6,2.0 , -2.0, \nand  2.2)  were  significantly  different  from  those  in  the  control  condition.  These \nhigher-order  interactions  arose  because  a  common-effect  schema  requires  only  one \nthe  presence  of  the  common  effect.  Figures  7b  presents \ncause  feature  to  explain \nthe  logarithm  of the  ratings  in  the  common-effect condition  for  those  test  exemplars \nin  which  the  common  effect is  present as  a  function  of the  number  of cause  features \npresent.  Ratings \nincreased \nmore  with  the  introduction \ncause \nof \nas \nsubsequent \ncompared \ncauses.  That  is,  participants \nconsidered  the  presence  of \nat \ncause \nexplaining  the  presence  of \nthe  common-effect \nto  be \nsufficient  grounds  to  grant \nan  exemplar  a \nrelatively \nhigh  category  membership \nrating  in  a  common-effect \ncategory. \ncontrast , \nFigure  7a  shows  a  linear \n\n4.5 \nbO  4.0 \n.= 'ill \n~ 3.5 \nOf) \n0 \n.....l  3.0 \n\n#  of Causes \n\n# of Effects \n\nObserved \n(CC Present) \n\nfirst \nto \n\n\u2022 \n\nObserved \n(CE  Present) \n\nFigure 4 \n\nleast \n\n2 \n\n3 \n\none \n\nPredicted \n\n2 \n\n3 \n\nthe \n\n2.5 \n\no \n\no \n\nIn \n\n\fincrease  in  (the  logarithm  of)  categorization  ratings  for  those  exemplars  in  which \nthe  common  cause  is  present  as  a  function  of  the  number  of effect  features.  In  the \npresence  of the  common  cause  each  additional  effect produced  a  constant increment \nto  log categorization ratings. \n\nFinally , Figure  3  also  indicates  that the  simple  feature  weights  differed  as  a function \nof  causal  schema.  In  the  common-cause  condition,  the  common-cause  (f1)  carried \ngreater  weight  than  the  three  effects  (f2,  f3 ,  f4).  In  contrast,  in  the  common-effect \ncondition  it  was  the  common-effect  (f4)  that  had  greater  weight  than  the  three \ncauses  (f1 ,  f2,  f3).  That  is ,  causal  networks  promote  the  importance  of  not  only \nspecific feature  combinations , but the importance of individual features  as  well. \n\nModel  Fitting \n\nTo  assess  whether  CMT  accounts  for  the  patterns  of  classification  found  in  this \nexperiment,  the  causal  models  of  Figure  1  were  fitted  to  the  category  membership \nratings  of  each  participant  in  the  common-cause  and  common-effect  conditions, \nrespectively. That is , the ratings  were predicted from the equation , \n\nRating (X) = K \u00a5 Likelihood (X;  c, m , b) \n\nwhere  Likelihood (X; c,  m ,  b)  is  the  likelihood  of exemplar X  as  a  function  of c,  m , \nand  b.  The  likelihood  equations  for  the  common-cause  and  common-effect  models \nshown  in  Table  1  were  used  for  common-cause  and  common-effect  participants , \nrespectively.  K  is  a  scaling  constant  that  brings  the  likelihood  into  the  range  0-100. \nFor  each  participant,  the  values  for  parameters  K ,  c,  m,  and  b  that  minimized  the \nsquared  deviation  between  the  predicted  and  observed  ratings  was  computed.  The \nbest  fitting  values  for  parameters  K ,  c,  m ,  and  b  averaged  over  participants  were \n846 ,  .578 ,  .214 ,  and  .437  in  the  common-cause  condition ,  and  876 ,  .522 ,  .325 ,  and \n.280  in  the  common-effect  condition.  The  predicted  ratings  for  each  exemplar  are \npresented  in  Table  1.  The  significantly  positive  estimate  for  m  in  both  conditions \nindicates  that  participants  categorization  performance  was  consistent  with  them \nassuming  the  presence  of  a  probabilistic  causal  mechanisms  linking  category \nfeatures.  Ratings  predicted  by  CMT  did  not  differ  from  observed  ratings  according \nto  chi-square tests:  )(\\16)=3.0 for  common cause, )(\\16)=5.3 for  common-effect. \n\nsensitivity \n\nthat  CMT  predicts  participants \n\nTo  demonstrate \nto  particular \ncombinations  of  features  when  categorizing ,  each  participants  predicted  ratings \nwere  subjected  to  the  same  regressions  that  were  performed  on  the  observed  ratings. \nThe  resulting  regression  weights  averaged  over  participants  are  presented  in  Figure \n3  superimposed on  the  weights  from  the  observed  data.  First,  Figure  3  indicates  that \nCMT  reproduces  participants  sensitivity  to  agreement  between  pairs  of  features \ndirectly  connected  by  causal  relationships  (f12 ,  f13 ,  f14  in  the  common-cause \ncondition ,  and  f14 ,  f24 ,  f34  in  the  common-effect  condition).  That  is ,  according  to \nboth  CMT  and  human  participants ,  category  membership  ratings  increase  when \npairs  of  features  confirm  causal  laws ,  and  decrease  when  they  violate  those  laws. \nSecond,  Figure  3  indicates  that  CMT  accounts  for  the  interactions  between  the \neffect features  in  the  common-cause  condition  (f12, f13 , f23)  and  also  for  the  higher(cid:173)\norder  feature  interactions  in  the  common-effect  condition  (f124 ,  f134,  f234 ,  f1234) , \nindicating  that  that  CMT  is  also  sensitive  to  the  asymmetries  inherent  in  causal \nrelationships.  The  predictions  of CMT  superimposed  on  the  observed  data  in  Figure \n4  confirm  that  CMT , like  the  human  participants , requires  only  one cause  feature  to \nthe  presence  of a  common  effect  (nonlinear  increase  in  ratings  in  Figure \nexplain \n4b)  whereas  CMT  predicts  a  linear increase in  log  ratings  as  one adds  effect features \nto  a  common  cause  (Figure  4a).  Finally ,  CMT  also  accounts  for  the  larger  weight \ngiven  to  the common cause and common-effect features  (Figure 3). \n\n\fDiscussion \n\nThe  current  results  support  CMTs  claims  that  people  have  a  representation  of  the \nprobabilistic  causal mechanisms  that link category  features,  and  that they  classify  by \nevaluating  whether  an  objects  combination  of  features  was  likely  to  have  been \ngenerated  by  those  mechanisms.  That  is , people  have  models  of the  world  that  lead \nthem  to  expect  a  certain  distribution  of features  in  category  members ,  and  consider \nexemplars good category members to  the extent they  manifest those expectations. \n\nOne  way  this  effect  manifested  itself  is  in  terms  of  the  importance  of  preserved \ncorrelations  between  features  directly  connected  by  causal  relationships.  An \nalternative  model  that  accounts  for  this  particular  result  assumes  that  the  feature \nspace  is  expanded  to  include  configural cues  encoding  the  confirmation  or  violation \nof  each  causal  relationship  [6].  However ,  such  a  model  treats  causal  links  as \nsymmetric  and  does  not consider interactions  among links.  As  a result, it does  not fit \nthe  common  effect data  as  well  as  CMT  (Figure  4b) ,  because  it is  unable  to  account \nfor  categorizers  sensitivity  to  the  higher-order  feature  interactions  that  emerge  as  a \nresult of causal asymmetries in  a complex network. \n\ntraditional  models  of  categorization  by  emphasizing  the \nCMT  diverges  from \nknowledge  people  possess  as  opposed  to  the  examples  they  observe.  Indeed ,  the \ncurrent  experiment  differed  from  many  categorization  studies  in  not  providing \nexamples  of  category  members.  As  a  result,  CMT  is  applicable  to  the  many  real(cid:173)\nworld  categories  about  which  people  know  far  more  than  they  have  observed  first \nhand  (e.g.,  scientific  concepts).  Of course, for  many  other  categories  people  observe \ncategory  members ,  and  the  nature  of  the  interactions  between  knowledge  and \nobservations  is  an  open  question  of  considerable  interest.  Using  the  same  materials \nas  in  the  current  study,  the  effects  of  knowledge  and  observations  have  been \northogonally  manipulated  with  the  finding  that  observations  had  little  effect  on \nclassification  performance  as  compared  to  the  theories  [7].  Thus , theories  may  often \ndominate categorization decisions even when observations are  available. \n\nAcknowledgments \n\nSupport  for \nthis  research  was  provided  by  funds  from  the  National  Science \nFoundation  (Grants  Number  SBR-98l6458  and  SBR  97-20304)  and  from  the \nNational Institute of Mental Health  (Grant Number ROl  MH58362). \n\nReferences \n\n[1]  Murphy,  G .  L. ,  &  Medin,  D .  L.  (1985).  The  role  of theories  in  conceptual  coherence. \nPsychological Review , 92, 289-316. \n\n[2]  Rehder,  B.  (1999).  A  causal  model  theory  of categorization .  In  Proceedin gs  of the  21st \nAnnual Meeting of the Cognitive Science Society  (pp. 595-600). Vancouver. \n\n[3]  Waldmann , M .R ., Holyoak , K.J .,  &  Fratianne, A.  (1995). Causal  models and  the  acquisition \nof category structure. Journal of Experimental Psychology:  General, 124 , 181-206 . \n\n[4]  Pearl,  J.  (1988).  Probabilistic  reasoning  in  intelligent  systems:  Networks  of plausible \ninference. San Mateo , CA: Morgan Kaufman. \n\n[5]  Salmon,  W.  C.  (1984).  Scientific  explanation  and  the  causa l  structure  of the  world. \nPrinceton , NJ:  Princeton University Press. \n\n[6]  Gluck, M.  A. ,  &  Bower, G.  H . (1988).  Evaluating  an  adaptive  network  model  of human \nlearning. Journal of Memory and Language, 27,166-195. \n\n[7]  Rehder, B., &  Hastie , R.  (2001). Causal knowledge and categories: The effects of causal  beliefs \non  categorization , induction,  and  similarity.  Journal  of Experimental  Psychology:  General,  130 , \n323-360. \n\n\f", "award": [], "sourceid": 1993, "authors": [{"given_name": "Bob", "family_name": "Rehder", "institution": null}]}