Performance in an Audiovisual Selective Attention Task Using Speech-Like Stimuli Depends on the Talker Identities, But Not Temporal Coherence.
Audiovisual integration of speech can benefit the listener by not only improving comprehension of what a talker is saying but also helping a listener select a particular talker's voice from a mixture of sounds. Binding, an early integration of auditory and visual streams that helps an observer alloc...
| Publicado en: | Trends in Hearing pp. 1 - 17 |
|---|---|
| Autores principales: | , , |
| Formato: | pictorial research tables/charts Journal Article |
| Publicado: |
Sage Publications Inc.
10/17/2023
|
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=173048843&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 173048843 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 23312165 HCEJ jtl: Trends in Hearing issn: 23312165 maglogo: N pubinfo: dt: 10/17/2023 pid: 344 pub: Sage Publications Inc. place: Thousand Oaks, California artinfo: ui: 173048843 173048843 173048843 10.1177/23312165231207235 173048843 ppf: 1 ppct: 16 formats: tig: atl: Performance in an Audiovisual Selective Attention Task Using Speech-Like Stimuli Depends on the Talker Identities, But Not Temporal Coherence. aug: au: Cappelloni, Madeline S. Mateo, Vincent S. Maddox, Ross K. affil: Biomedical Engineering, 6927University of Rochester, Rochester, NY, USA sug: subj: Audiovisuals Evaluation Selective Attention Speech Perception Auditory Perception Task Performance and Analysis Voice Recognition Systems Utilization Human Listening Effort Sensory Stimulation Cues Speech Articulation Tests Perceptual Distortion Descriptive Statistics Confidence Intervals Funding Source ab: Audiovisual integration of speech can benefit the listener by not only improving comprehension of what a talker is saying but also helping a listener select a particular talker's voice from a mixture of sounds. Binding, an early integration of auditory and visual streams that helps an observer allocate attention to a combined audiovisual object, is likely involved in processing audiovisual speech. Although temporal coherence of stimulus features across sensory modalities has been implicated as an important cue for non-speech stimuli (Maddox et al., 2015), the specific cues that drive binding in speech are not fully understood due to the challenges of studying binding in natural stimuli. Here we used speech-like artificial stimuli that allowed us to isolate three potential contributors to binding: temporal coherence (are the face and the voice changing synchronously?), articulatory correspondence (do visual faces represent the correct phones?), and talker congruence (do the face and voice come from the same person?). In a trio of experiments, we examined the relative contributions of each of these cues. Normal hearing listeners performed a dual task in which they were instructed to respond to events in a target auditory stream while ignoring events in a distractor auditory stream (auditory discrimination) and detecting flashes in a visual stream (visual detection). We found that viewing the face of a talker who matched the attended voice (i.e., talker congruence) offered a performance benefit. We found no effect of temporal coherence on performance in this task, prompting an important recontextualization of previous findings. pubtype: Academic Journal doctype: pictorial research tables/charts Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|