CHI '95 ProceedingsTopIndexes
PostersTOC

A comparison of speech and mouse/keyboard GUI navigation.

Ron Van Buskirk, *Mary LaLomia

IBM Corporation, 1217
1000 NW 51st Street
Boca Raton, FL 33432
revanbus@bocaraton.ibm.com
mlalomia@vnet.ibm.com

© ACM

Abstract

We compared two speaker-independent, navigation systems (discrete and continuous) on 11 tasks, measuring accuracy, perceived performance, task time, and perceived system usability. Ten IBM and temporary help agency employees with GUI experience participated. Their ages ranged from 25 to 55 years. The participants completed 11 tasks on both systems using voice or keyboard. The participants began the set of tasks on a randomly selected navigator, filled out a questionnaire about the perceived system speed and accuracy, completed the same tasks using the keyboard, then repeated the same procedure on a second system and keyboard. The voice navigator tasks took approximately twice as long as the keyboard tasks. Additionally, the survey results showed that participants' acceptance of the system was quite sensitive to small changes in system response time. The slowest tasks were the ones with precise cursor or window movement, the fastest were ones only requiring brief commands. The results are discussed in terms of recommendations for designing speech into GUIs.

Keywords:

Speech navigation, continuous speech recognition, discrete speech recognition.

Introduction

As the price of computer hardware decreases and speech recognition technology matures, speech input will become more common. Speech input appears most useful for tasks that require small vocabularies, which users perform while their hands and/or eyes are busy. Consequently, voice navigators, which allow the user to control their computer with speech commands, are a promising voice application.

The main objective of this study was to evaluate two voice navigators using two different technologies and to compare speech navigation with keyboard navigation. One is a discrete speech recognizor, where the user issues commands separated with a brief pause. The other is a continuous speech recognition system that allows the user to issue several commands at once, without pausing.

Methodology

Participants performed 11 tasks with each navigator and the keyboard to evaluate the speed, accuracy and usability. The tasks included six standard system-management tasks (e.g., finding files, moving windows) and five hands-busy, eyes-busy tasks (e.g., data entry, italicizing text). Afterward, they completed rating scales of usability features.

Results and Discussion

The keyboard tasks took approximately half the time as the navigator tasks. In addition, very small differences in response time affected the participants' acceptance of the navigators' speed. An error analysis on the misrecognitions pointed to several improvements in the navigators' vocabularies.

There appeared to be little difference between the continuous and discrete recognizors, with little difference between task times and accuracy. However, because the continuous navigator had to process several commands at once, it had a slower response time. Overall task time was not affected, but the rated acceptability of its speed was much lower due to its increased response time.

The best tasks for speech input were tasks in which the user has to issue brief commands using a small vocabulary. The least suitable were tasks in which the speaker had to precisely control a cursor or issue a command from a large vocabulary.

Also we identified several important factors affecting navigation use, including system response time and experience with voice technology, which will be the basis of future study.