Learning and choosing in an uncertain world: An investigation of the explore–exploit dilemma in static and dynamic environments.
How do people solve the explore–exploit trade-off in a changing environment? In this paper we present experimental evidence from an “observe or bet” task, in which people have to determine when to engage in information-seeking behavior and when to switch to reward-taking actions. In particular we fo...
| Publicado en: | Cognitive Psychology Vol. 85; pp. 43 - 78 |
|---|---|
| Autores principales: | , , |
| Formato: | Artículo |
| Publicado: |
Academic Press Inc.
Mar2016
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ssf&AN=112827193&site=ehost-live header: @attributes: shortDbName: ssf uiTerm: 112827193 longDbName: Social Sciences Full Text (H.W. Wilson) uiTag: AN controlInfo: bkinfo: jinfo: jid: 00100285 COP jtl: Cognitive Psychology issn: 00100285 maglogo: N pubinfo: dt: Mar2016 vid: 85 pid: 735 pub: Academic Press Inc. artinfo: ui: 112827193 10.1016/j.cogpsych.2016.01.001 ppf: 43 ppct: 35 formats: tig: atl: Learning and choosing in an uncertain world: An investigation of the explore–exploit dilemma in static and dynamic environments. aug: au: Navarro, Daniel J. Newell, Ben R. Schulze, Christin affil: School of Psychology, University of Adelaide, Australia School of Psychology, University of New South Wales, Australia su: Information-seeking behavior Human behavior Mathematical models Learning problems sug: subj: Information-seeking behavior Human behavior Mathematical models Learning problems keyword: Decision making Decisions from experience Dynamic environments Explore–exploit dilemma Decision making Decisions from experience Dynamic environments Explore–exploit dilemma ab: How do people solve the explore–exploit trade-off in a changing environment? In this paper we present experimental evidence from an “observe or bet” task, in which people have to determine when to engage in information-seeking behavior and when to switch to reward-taking actions. In particular we focus on the comparison between people’s behavior in a changing environment and their behavior in an unchanging one. Our experimental work is motivated by rational analysis of the problem that makes strong predictions about information search and reward seeking in static and changeable environments. Our results show a striking agreement between human behavior and the optimal policy, but also highlight a number of systematic differences. In particular, we find that while people often employ suboptimal strategies the first time they encounter the learning problem, most people are able to approximate the correct strategy after minimal experience. In order to describe both the manner in which people’s choices are similar to but slightly different from an optimal standard, we introduce four process models for the observe or bet task and evaluate them as potential theories of human behavior. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: N holdings: @attributes: islocal: N |
|---|