{"title": "A Bayesian Model Predicts Human Parse Preference and Reading Times in Sentence Processing", "book": "Advances in Neural Information Processing Systems", "page_first": 59, "page_last": 65, "abstract": null, "full_text": "A Bayesian Model Predicts Human Parse\nPreference and Reading Times in Sentence\n\nProcessing\n\nSrini Narayanan\n\nDaniel Jurafsky\n\nSRI International and ICSI Berkeley University of Colorado, Boulder\n\nsnarayan@cs.berkeley.edu\n\njurafsky@colorado.edu\n\nAbstract\n\nNarayanan and Jurafsky (1998) proposed that human language compre-\nhension can be modeled by treating human comprehenders as Bayesian\nreasoners, and modeling the comprehension process with Bayesian de-\ncision trees. In this paper we extend the Narayanan and Jurafsky model\nto make further predictions about reading time given the probability of\ndifference parses or interpretations, and test the model against reading\ntime data from a psycholinguistic experiment.\n\n1\n\nIntroduction\n\nNarayanan and Jurafsky (1998) proposed that human language comprehension can be mod-\neled by treating human comprehenders as Bayesian reasoners, and modeling the compre-\nhension process with Bayesian decision trees. In this paper, we show that the model ac-\ncounts for parse-preference and reading time data from a psycholinguistic experiment on\nreading time in ambiguous sentences.\n\nParsing, (generally called \u2018sentence processing\u2019 when we are referring to human parsing),\nis the process of building up syntactic interpretations for a sentence from an input sequence\nof written or spoken words. Ambiguity is extremely common in parsing problems, and\nprevious research on human parsing has focused on showing that many factors play a role\nin choosing among the possible interpretations of an ambiguous sentence.\n\nWe will focus in this paper on a syntactic ambiguity phenomenon which has been repeat-\nedly investigated: the main-verb (MV), reduced relative (RR) local ambiguity (Frazier &\nRayner, 1987; MacDonald, Pearlmutter, & Seidenberg, 1994; McRae, Spivey-Knowlton,\n& Tanenhaus, 1998, inter alia) In this ambiguity, a pre\ufb01x beginning with a noun-phrase and\nan ambiguous verb-form could either be continued as a main clause (as in 1a), or turn out\nto be a relative clause modi\ufb01er of the \ufb01rst noun phrase (as in 1b).\n\n1. a. The cop arrested the forger.\n\nb. The cop arrested by the detective was guilty of taking bribes.\n\nMany factors are known to in\ufb02uence human parse preferences. One such factor is the dif-\nferent lexical/morphological frequencies of the simple past and participial forms of the am-\nbiguous verbform (arrested, in this case). Trueswell (1996) found that verbs like searched,\nwith a frequency-based preference for the simple past form, caused readers to prefer the\n\n\fmain clause interpretation, while verbs like selected, had a preference for a participle read-\ning, and supported the reduced relative interpretation.\n\nThe transitivity preference of the verb also plays a role in human syntactic disambiguation.\nSome verbs are preferably transitive, where others are preferably intransitive. The reduced\nrelative interpretation, since it involves a passive structure, requires that the verb be tran-\nsitive. MacDonald, Pearlmutter, and Seidenberg (1994), Trueswell, Tanenhaus, and Kello\n(1994) and other have shown that verbs which are biased toward an intransitive interpreta-\ntion also bias readers toward a main clause interpretation.\n\nPrevious work has shown that a competition-integration model developed by Spivey-\nKnowlton (1996) could model human parse preference in reading ambiguous sentences\n(McRae et al., 1998). While this model does a nice job of accounting for the reading-time\ndata, it and similar \u2018constraint-based\u2019 models rely on a complex set of feature values and\nfactor weights which must be set by hand. Narayanan and Jurafsky (1998) proposed an\nalternative Bayesian approach for this constraint-combination problem. A Bayesian ap-\nproach offers a well-understood formalism for de\ufb01ning probabilistic weights, as well as for\ncombining those weights. Their Bayesian model is based on the probabilistic beam-search\nof Jurafsky (1996), in which each interpretation receives a probability, and interpretations\nwere pruned if they were much worse than the best interpretation. The model predicted\nlarge increases in reading time when unexpected words appeared which were only compat-\nible with a previously-pruned interpretation. The model was thus only able to characterize\nvery gross timing effects caused by pruning of interpretations.\n\nIn this paper we extend this model\u2019s predictions about reading time to other cases where\nthe best interpretation turns out to be incompatible with incoming words. In particular, we\nsuggest that any evidence which causes the probability of the best interpretation to drop\nbelow its next competitor will also cause increases in reading time.\n\n2 The Experimental Data\n\nWe test our model on the reading time data from McRae et al. (1998), an experiment fo-\ncusing on the effect of thematic \ufb01t on syntactic ambiguity resolution. The thematic role of\nnoun phrase \u201cthe cop\u201d in the pre\ufb01x \u201cThe cop arrested\u201d is ambiguous. In the continuation\n\u201dThe cop arrested the crook\u201d, the cop is the agent. In the continuation \u201cThe cop arrested by\nthe FBI agent was convicted for smuggling drugs\u201d, the cop is the theme. The probabilistic\nrelationship between the noun and the head verb (\u201carrested\u201d) biases the thematic disam-\nbiguation decision. For example, \u201ccop\u201d is a more likely agent for \u201carrest\u201d, while \u201ccrook\u201d is\na more likely theme. McRae et al. (1998) showed that this \u2018thematic \ufb01t\u2019 between the noun\nand verb affected phrase-by-phrase reading times in sentences like the following:\n\n2. a. The cop / arrested by / the detective / was guilty / of taking / bribes.\n\nb. The crook / arrested by / the detective / was guilty / of taking / bribes.\n\nIn a series of experiment on 40 verbs, they found that sentences with good agents (like cop\nin 2a) caused longer reading times for the phrase the detective than sentences with good\nthemes (like crook in 2b). Figure 1 shows that at the initial noun phrase, reading time is\nlower for good-agent sentences than good-patient sentences. But at the NP after the word\n\u201cby\u201d, reading time is lower for good-patient sentences than good-agent sentences.1\n\n1In order to control for other in\ufb02uences on timing, McRae et al. (1998) actually report reading\ntime deltas between a reduced relative and non-reduced relative for. It is these deltas, rather than raw\nreading times, that our model attempts to predict.\n\n\f \n\ni\n\no\nt\n \nd\ne\nr\na\np\nm\no\nc\n(\n \ns\ne\nm\nT\ng\nn\nd\na\ne\nR\nd\ne\ns\na\ne\nr\nc\nn\n\n \n\n \n\ni\n\nI\n\n)\nl\no\nr\nt\nn\no\nc\n\n7 0\n\n6 0\n\n5 0\n\n4 0\n\n3 0\n\n2 0\n\n1 0\n\n0\n\nThe cop/crook arrested by\n\n the detective\n\nGood Agent\n\nGood Patient\n\nFigure 1: Self-paced reading times (from Figure 6 of McRae et al. (1998))\n\nAfter introducing our model in the next section, we show that it predicts this cross-over in\nreading time; longer reading time for the initial NP in good-patient sentences, but shorter\nreading time for the post-\u201cby\u201d NP in good-patient sentences.\n\n3 The Model and the Input Probabilities\n\nIn the Narayanan and Jurafsky (1998) model of sentence processing, each interpretation of\nan ambiguous sentence is maintained in parallel, and associated with a probability which\ncan be computed via a Bayesian belief net. The model pruned low-probability parses, and\nhence predicted increases in reading time when reading a word which did not \ufb01t into any\navailable parse. The current paper extends the Narayanan and Jurafsky (1998) model\u2019s pre-\ndictions about reading time. The model now also predicts extended reading time whenever\nan input word causes the best interpretation to drop in probability enough to switch in rank\nwith another interpretation.\n\nThe model consists of a set of probabilities expressing constraints on sentence processing,\nand a network that represents their independence relations:\n\nSource\nData\nP(Agent  verb, initial NP)\nMcRae et al. (1998)\nP(Patient  verb, initial NP)\nMcRae et al. (1998)\nP(Participle  verb)\nBritish National Corpus counts\nP(SimplePast verb)\nBritish National Corpus counts\nP(transitive  verb)\nTASA corpus counts\nP(intransitive  verb)\nTASA corpus counts\nP(RR  initial NP, verb-ed, by)\nMcRae et al. (1998) (.8, .2)\nP(RR  initial NP, verb-ed, by,the)\nMcRae et al. (1998) (.875. .125)\nP(Agent  initial NP, verb-ed, by, the, NP) McRae et al. (1998) (4.6 average)\nSCFG counts from Penn Treebank\nP(MC  SCFG pre\ufb01x)\nP(RR  SCFG pre\ufb01x)\nSCFG counts from Penn Treebank\n\nThe \ufb01rst constraint expresses the probability that the word \u201ccop\u201d, for example, is an agent,\ngiven that the verb is \u201carrested\u201d. The second constraint expresses the probability that it is\na patient. The third and fourth constraints express the probability that the \u201d-ed\u201d form of\nthe verb is a participle versus a simple past form (for example P(Participle  \u201carrest\u201d)=.81).\nThese were computed from the POS-tagged British National Corpus. Verb transitivity\nprobabilities were computed by hand-labeling subcategorization of 100 examples of each\nverb in the TASA corpus. (for example P(transitive  \u201centertain\u201d)=.86). Main clause prior\nprobabilities were computed by using an SCFG with rule probabilities trained on the Penn\n\n\fTreebank version of the Brown corpus. See Narayanan and Jurafsky (1998) and Jurafsky\n(1996) for more details on probability computation.\n\n4 Construction Processing via Bayes nets\n\nUsing Belief nets to model human sentence processing allows us to a) quantitatively eval-\nuate the impact of different independence assumptions in a uniform framework, b) directly\nmodel the impact of highly structured linguistic knowledge sources with local conditional\nprobability tables, while well known algorithms for updating the Belief net (Jensen (1995))\ncan compute the global impact of new evidence, and c) develop an on-line interpretation\nalgorithm, where partial input corresponds to partial evidence on the network, and the up-\ndate algorithm appropriately marginalizes over unobserved nodes. So as evidence comes\nin incrementally, different nodes are instantiated and the posterior probability of different\ninterpretations changes appropriately.\n\nlatent variables that render top-down ( \n\nThe crucial insight of our Belief net model is to view speci\ufb01c interpretations as values of\n) conditionally\nindependent (d-separate them (Pearl, 1988)). Thus syntactic, lexical, argument structure,\nand other contextual information acts as prior or causal support for an interpretation, while\nbottom-up phonological or graphological and other perceptual information acts as likeli-\nhood, evidential, or diagnostic support.\n\n) and bottom-up evidence (\u0003\u0002\n\nTo apply our model to on-line disambiguation, we assume that there are a set of interpre-\ntations ( \u0004\u0006\u0005\b\u0007\n\t\f\u000b\n\u000b\f\u000b\b\u0005\u000e\r\u0010\u000f\u0012\u0011\u0014\u0013\n) that are consistent with the input data. At different stages of the\ninput, we compute the posterior probabilities of the different interpretations given the top\ndown and bottom-up evidence seen so far. 2\n\nV = examine-ed type_of(Subj) = witness\n\nP(Arg|v)\n\nP(Tense|v)\n\nP(A | v, ty(Subj))\n\nP(T | v, ty(Subj))\n\nArg\n\nTense\n\nSemantic_fit\n\nAND\n\nAND\n\nTense = past\nSem_fit = Agent\n\nMV\n\nthm\n\nArg = trans\nTense = pp\nSem_fit = Theme\n\nRR\n\nthm\n\nFigure 2: The Belief net that represents lexical and thematic support for the two interpreta-\ntions.\n\nFigure 2 reintroduces the basic structure of our belief net model from Narayanan and Ju-\nrafsky (1998). Our model requires conditional probability distributions specifying the pref-\nerence of every verb for different argument structures, as well its preference for different\ntenses. We also compute the semantic \ufb01t between possible \ufb01llers in the input and different\ninterpre-\ntations require the conjunction of speci\ufb01c values corresponding to tense, semantic \ufb01t and\ninterpretation requires the transitive\n\nconceptual roles of a given predicate. As shown in Figure 2, the \u0015\u0017\u0016\nargument structure features. Note that only the \u0018\u0019\u0018\n\nargument structure.\n\nand \u0018\u0019\u0018\n\n2In this paper, we will focus on the support from thematic, and syntactic features for the Re-\nduced Relative (RR) and Main Verb (MV) interpretations at different stages of the input for the\n\nexamples we saw earlier. So we will have two interpretations \u001a\u001c\u001b\u001e\u001d\b\u001a \u001f\"!$# where %\u0012&\u0006\u001a'\u001b)(\n2$3\n\n.\n\n*\f+,\u001d-*).,/10\n\n\u001d4%\u0012&\u0006\u001a \u001f)(\n\n\u001d-*\n\n/,06575\n\n\u0001\n*\n+\n.\n\fS\n\n[.48] S\u2212> NP [V ...\n\nNP\n\n#1[]\n\nVP\n\nV\n\nthe witness examined\n\nVP\n\n \n\nMAIN VERB\n\n[.14] NP\u2212> NP XP\n\nS\n\n[.92] S\u2212> NP ...\n\nNP\n\nNP\n\nVP\n\nThe witness\n\nexamined\n\nREDUCED RELATIVE\n\nFigure 3: The partial syntactic parse trees for the \u0015\u0017\u0016\ning an \u0002\u0001\u0004\u0003\u0004\u0005\n\ngenerating grammar.\n\nand the \u0018\u0019\u0018\n\ninterpretations assum-\n\nMAIN CLAUSE\n\nS\n\nNP\n\nVP\n\nDet\n\nN\n\n V\n\nXP\n\nREDUCED RELATIVE\n\nS\n\nNP\n\nXP\n\nNP\n\nVP\n\nDet\n\nN\n\nThe witness examined\n\n The witness examined\n\nFigure 4: The Bayes nets for the partial syntactic parse trees\n\n\u0004\b\u0007\n\nparent and child nodes, d) the \n\nThe conditional probability of a construction given top-down syntactic evidence \u0006\nthe \u0002\u0001\u0004\u0003\u0004\u0005\n\nis relatively simple to compute in an augmented-stochastic-context-free formalism (partial\nparse trees shown in Figure 3 and the corresponding bayes net in Figure 4). Recall that\nprior probability gives the conditional probability of the right hand side of a\nrule given the left hand side. The Inside/Outside algorithm applied to a \ufb01xed parse tree\nstructure is obtained exactly by casting parsing as a special instance of belief propagation.\nThe correspondences are straightforward a) the parse tree is interpreted as a belief network.\nb)the non-terminal nodes correspond to random variables, the range of the variables being\nthe non-terminal alphabet, c) the grammar rules de\ufb01ne the conditional probabilities linking\nnonterminal at the root, as well as the terminals at the\nleaves represent conditioning evidence to the network, and e) Conditioning on this evidence\nproduces exactly the conditional probabilities for each nonterminal node in the parse tree\nand the joint probability distribution of the parse. 3\nThe overall posterior ratio requires propagating the conjunctive impact of syntactic and\nlexical/thematic sources on our model. Furthermore, in computing the conjunctive im-\n, we use the\nNOISY-AND (assumes exception independence) model (Pearl, 1988) for combining con-\ninterpretations. At various points, we\ncompute the posterior support for the different interpretations using the following equa-\n. The \ufb01rst term is\ntion.\n\npact of the lexical/thematic and syntactic support to compute \u0015\u0017\u0016\njunctive sources. In the case of the \u0018\u0019\u0018\n\u000f\u0019\n\n\t\u0014\u0013\u0016\u0015\u0018\u0017\n\nand \u0018\u0019\u0018\n\n\t\u001e\u001d\u001f\u0015\u0018\u0017\n\n\u000f\u001b\u001a\n\n\u000f\u000b\n\n\t\u0010\u000f\n\n\u0004\b\t\n\n\u0004\b\t\n\n\u0004\b\t\n\nand \u0015\u0017\u0016\n\u0004\b\t\n\u0002\f\u000e\n\n\u0012\u0011\n\n\u0002\f\u000e\r\n3One complication is that the the conditional distribution in a parse tree %\u0012&! ,\u001d#\"\u0012(\nthe product distribution %\u0012&! \n/-%\u0012&!\"\u0012(\nsible to generalize the belief propagation equations to admit conjunctive distributions %\u0012&! ,\u001d#\"\u0012(\nand %\u0012&\b$\"\u001d\n&\b2\n/1&\n/-%\u0012&\b0\n\u001d12\nand the causal support becomes 3\u0010&\b'\n\n. The diagnostic (inside) support becomes &\n&\b'\n&!=)/-%\u0012&\b'\n\nis not\n(it is the conjunctive distribution). However, it is pos-\n\nhttp://www.icsi.berkeley.edu/ snarayan/scfg.ps).\n\n0)(+*-,\n<\u0003/\n\n(details can be found at\n\n0546(879,\n\n:;3\u0010&!<\n\n./&\n\n\u001d#=\n\n/1&\n\n&\b0\n\n\u0012\u0011\n\n\n\n\u000f\n\u0006\n\u0006\n\n\u0006\n\n\u0006\n\n\u001c\n\u000f\n$\n/\n(\n$\n$\n/\n$\n/\n3\n(\n%\n/\n/\n(\n'\n/\n/\n(\n\fthe syntactic support while the second is the lexical and thematic support for a particular\ninterpretation ( \t\n\u0018\u0019\u0018\n5 Model results\n\n\u0015\u0017\u0016\n\n).\n\nMV/RR\n\n9\n\n8\n\n7\n\n6\n\n5\n\n4\n\n3\n\n2\n\n1\n\n0\n\nNP verbed\n\n by\n\n the\n\n NP\n\nModel Good Agent Human Good Agent Model Good Patient Human Good Patient\n\nFigure 5: Completion data\n\nWe tested our model on sentences with the different verbs in McRae et al. (1998). For\neach verb, we ran our model on sentences with Good Agents (GA) and Good Patients (GP)\nfor the initial NP. Our model results are consistent with the on-line disambiguation studies\nwith human subjects (human performance data from McRae et al. (1998)) and show that a\nBayesian implementation of probabilistic evidence combination accounts for garden-path\ndisambiguation effects.\n\nFigure 5 shows the \ufb01rst result that pertains to the model predictions of how thematic \ufb01t\nmight in\ufb02uence sentence completion times. Our model shows close correspondence to the\nhuman judgements about whether a speci\ufb01c ambiguous verb was used in the Main Clause\n(MV) or reduced relative (RR) constructions. The human and model predictions were con-\nducted at the verb(The crook arrested), by (the crook arrested by), the (the crook arrested\nby the) and Agent NP (the crook arrested by the detective). As in McRae et al. (1998)\nthe data shows that thematic \ufb01t clearly in\ufb02uenced the gated sentence completion task. The\nprobabilistic account further captured the fact that at the by phrase, the posterior proba-\nbility of producing an RR interpretation increased sharply, thematic \ufb01t and other factors\n\n)\n\nR\nR\nP\n\n(\n\n/\n)\n\nC\nM\nP\n\n(\n\n2.5\n\n2\n\n1.5\n\n1\n\n0.5\n\n0\n\n2.1\n\n0.7\n\n0.541\n\n0.13\n\nThe crook/detective\n\n the\n\narrested by\n\n0.1\n0.04\n\n detective\n\n)\n\nX\n\n(\n\nP\n\n0.9\n\n0.8\n\n0.7\n\n0.6\n\n0.5\n\n0.4\n\n0.3\n\n0.2\n\n0.1\n\n0\n\nThe cop arrested by\n\n the detective\n\n(a)\n\nGood Agent initial NP\n\nGood Patient Initial NP\n\nGood Agent Main Clause\n\nGood Agent RR\n\n(b)\n\nFigure 6: a) MV/RR for the ambiguous region showing a \ufb02ip for the Good Agent (ga) case.\nb) P(MV) and P(RR) for the Good Patient and Good Agent cases.\n\n\u0011\n\t\n\fin\ufb02uenced both the sharpness and the magnitude of the increase.\n\nThe second result pertains to on-line reading times. Figure 6 shows how the human reading\ntime reduction effects (reduced compared to unreduced interpretations) increase for Good\nAgents (GA) but decrease for Good Patients in the ambiguous region. This explains the\nreading time data in Figure 1. Our model predicts this larger effect from the fact that the\nmost probable interpretation for the Good Agent case \ufb02ipsfrom the MV to the RR interpre-\ntation in this region. No such \ufb02ip occurs for the Good Patient (GP) case. In Figure 6(a), we\nsee that the GP results already have the MV/RR ratio less than one (the RR interpretation\nis superior) while a \ufb02ip occurs for the GA sentences (from the initial state where MV/RR\n\u0001 ). Figure 6 (b) shows a more detailed view of\n\u0002\u0001\nthe GA sentences showing the crossing point where the \ufb02ipoccurs. This \ufb01nding is fairly\nrobust (\n\nof GA examples) and directly predicts reading time dif\ufb01culties.\n\nto the \ufb01nal state where MV/RR\n\n\u0004\u0006\u0005\b\u0007\n\n6 Conclusion\n\nWe have shown that a Bayesian model of human sentence processing is capable of modeling\nreading time data from a syntactic disambiguation task. A Bayesian model extends current\nconstraint-satisfaction models of sentence processing with a principled way to weight and\ncombine evidence. Bayesian models have not been widely applied in psycholinguistics. To\nour knowledge, this is the \ufb01rst study showing a direct correspondence between the time\ncourse of maintaining the best a posteriori interpretation and reading time dif\ufb01culty.\n\nWe are currently exploring how our results on \ufb02ippingof preferred interpretation could\nbe combined with Hale (2001)\u2019s proposal that reading time correlates with surprise(a sur-\nprising (low probability) word leads to large amounts of probability mass to be pruned) to\narrive at a structured probabilistic account of a wide variety of psycholinguistic data.\n\nReferences\n\nFrazier, L., & Rayner, K. (1987). Resolution of syntactic category ambiguities: Eye movements in\n\nparsing lexically ambiguous sentences. Journal of Memory and Language, 26, 505\u2013526.\n\nHale, J. (2001). A probabilistic earley parser as a psycholinguistic model. Proceedings of NAACL-\n\n2001.\n\nJensen, F. (1995). Bayesian Networks. Springer-Verlag.\nJurafsky, D. (1996). A probabilistic model of lexical and syntactic access and disambiguation. Cog-\n\nnitive Science, 20, 137\u2013194.\n\nMacDonald, M.C., Pearlmutter, N.J., & Seidenberg, M.(1994). The lexical nature of syntactic ambi-\n\nguity resolution. Psychological Review, 101, 676-703.\n\nMcRae, K., Spivey-Knowlton, M., & Tanenhaus, M. K.(1998). Modeling the effect of thematic \ufb01t\n(and other constraints) in on-line sentence comprehension. Journal of Memory and Language,38,\n283\u2013312.\n\nNarayanan, S., & Jurafsky, D. (1998). Bayesian models of human sentence processing. In COGSCI-\n\n98, pp. 752\u2013757 Madison, WI. Lawrence Erlbaum.\n\nPearl, J. (1988). Probabilistic Reasoning in Intelligent Systems: Networks of Plausible Inference.\n\nMorgan Kaufman, San Mateo, Ca.\n\nSpivey-Knowlton, M.(1996). Integration of visual and linguistic information: Human data and model\n\nsimulations. Ph.D. Thesis, University of Rochester, 1996.\n\nTrueswell, J. C.(1996). The role of lexical frequency in syntactic ambiguity resolution. Journal of\n\nMemory and Language, 35, 566-585.\n\nTrueswell, J. C., Tanenhaus, M. K., & Kello, C. (1994). Verb-speci\ufb01c constraints in sentence pro-\ncessing: Separating effects of lexical preference from garden-paths. Journal of Experimental\nPyschology: Learning, Memory and Cognition, 19(3), 528\u2013553.\n\n\u0003\n\f", "award": [], "sourceid": 2130, "authors": [{"given_name": "S.", "family_name": "Narayanan", "institution": null}, {"given_name": "Daniel", "family_name": "Jurafsky", "institution": null}]}