FIELD NOTE / 2026.09.136 MIN READ / 5 SOURCES

Card, Moran, and Newell’s GOMS: Modeling the Human Side of Computer Performance

GOMS turned skilled interaction into a hierarchy of goals, operators, methods, and selection rules, giving HCI researchers a way to predict expert task performance before building or testing an interface.

HCI tried to turn psychology into engineering

By the late 1970s, interactive computing was creating a new engineering problem. Designers could measure processor speed and memory capacity precisely, but the human operating the system was often described with intuition rather than predictive models. Stuart Card, Thomas Moran, and Allen Newell at Xerox PARC and Carnegie Mellon sought a more rigorous approach. Their work treated skilled computer use as information processing that could be decomposed, timed, and compared. In a 1980 analysis of text editing, they described user behavior in terms of goals, operators, methods, and selection rules—the elements that became known collectively as GOMS.[1] This did not attempt to model every emotion or discovery a person might experience. It focused on routine, learned behavior where users know what they want to accomplish. That narrower scope made the model useful as an engineering instrument rather than a theory of all human thought.

Predictive models targeted skilled routine behavior

GOMS is strongest when a task has recognizable procedures and an experienced user can execute them without prolonged problem solving. The framework deliberately trades psychological completeness for enough structure to estimate how an interface supports familiar work.

GOMS decomposed action into goals, operators, methods, and rules

The four parts of GOMS describe different layers of task performance. Goals state what the user is trying to achieve. Operators are elementary perceptual, cognitive, or motor actions available to the user. Methods are procedures composed from those operators, and selection rules specify how a user chooses among alternative methods when more than one can achieve the same goal. Card, Moran, and Newell used this structure to analyze text-editing behavior and reported that method choices could be predicted with substantial accuracy in their studied tasks.[1] The decomposition gave designers a vocabulary for locating inefficiency. A slow task might contain too many operators, require a cumbersome method, or force users to choose between inconsistent procedures. Rather than saying an interface “feels complicated,” an analyst could describe where the complication enters the user’s procedure.

Selection rules made alternatives part of the model

Users often know several ways to perform the same command: keyboard shortcut, menu, mouse gesture, or command sequence. GOMS made that plurality explicit instead of pretending that one canonical procedure defines the interface.

The text-editing studies made the framework empirical

GOMS grew out of attempts to explain real editing behavior rather than purely formal speculation. The 1980 text-editing paper connected observable user actions to an information-processing description and tested whether users’ choices among methods could be predicted.[1] That empirical grounding mattered because an analytic model is useful only if its decomposition corresponds reasonably well to what skilled users actually do. The researchers could compare predicted procedural behavior with observed editing sequences and refine assumptions about mental preparation, command selection, and execution. This approach also helped distinguish learning from execution. GOMS was not mainly a model of how novices discover a command; it was a model of how practiced users carry out known tasks. That distinction became central to later HCI: an interface can be easy to learn yet inefficient for experts, or efficient for experts while difficult for novices to understand.

A model can be useful without simulating the whole mind

Card, Moran, and Newell’s engineering approach succeeded by modeling only the aspects of cognition and motor behavior needed for a particular prediction. That principle became a recurring strategy in quantitative HCI.

The Keystroke-Level Model made prediction practical

The Keystroke-Level Model, introduced by Card, Moran, and Newell in 1980, distilled the broader modeling program into a particularly practical technique for predicting the execution time of routine interactive tasks.[2] An analyst represents a procedure as a sequence of operators such as keystrokes, pointing actions, homing between devices, system response waits, and mental preparation. Standard time estimates are then combined to predict how long an expert user should take under specified conditions. The attraction was not merely accuracy; it was cost. Designers could compare two command designs on paper without recruiting users for every tiny decision. If one method required extra homing movements or more mental preparation, the model could expose that burden before implementation. KLM therefore turned human performance into something that could participate in early design tradeoffs alongside code size, hardware cost, and response time.

Paper analysis could precede expensive prototypes

Because KLM works from an explicit task sequence, it can be applied before a full interface exists. That made predictive human-performance modeling especially attractive in projects where prototypes or field tests were costly.

Command Language Grammar connected tasks to interface structure

Thomas Moran extended the broader engineering program with the Command Language Grammar, which analyzed command interfaces across multiple descriptive levels. His 1981 paper distinguished task, semantic, syntactic, and interaction structures, providing a way to relate what users want to accomplish to how commands are expressed.[4] The work complemented GOMS by showing that interface complexity can appear at different boundaries. Two systems may support the same task semantics while requiring very different command syntax, or they may expose inconsistent interaction patterns that increase what users must remember. This was especially relevant in an era of text commands and menu systems, but the lesson survives in graphical and touch interfaces: interaction design is a mapping between user intentions and available operations. A good model helps identify where that mapping becomes unnecessarily indirect.

Project Ernestine showed that predictive models could matter economically

A famous later demonstration came from Project Ernestine, in which researchers used GOMS-based analysis to evaluate a proposed workstation for telephone operators. The model predicted that the new design would actually be slower for key tasks, despite assumptions that newer technology should improve productivity. Field data supported the prediction, and the consequences were financially significant because tiny differences repeated across many operators and calls could accumulate into large annual costs.[5] The episode became an important case study because it showed where analytic HCI can be unusually powerful: high-volume routine work. A difference of fractions of a second may seem trivial in one interaction but become substantial when multiplied by millions of repetitions. Predictive models therefore offered management-relevant evidence, not merely academic description.

The Psychology of Human-Computer Interaction unified a research program

Card, Moran, and Newell consolidated these ideas in their 1983 book The Psychology of Human-Computer Interaction. The book presented a broader model-human-processor framework and treated interaction as a domain in which perceptual, cognitive, and motor performance could be connected to system design.[3] Its influence came partly from offering HCI a recognizable engineering identity. Instead of evaluating interfaces only after implementation, designers could reason prospectively about task structure and user performance. The framework also encouraged explicit assumptions: which users are modeled, what they know, how long basic operations take, and where system delays enter. Critics later challenged the field to account for situated action, emotion, collaboration, and other phenomena that GOMS does not capture, but those critiques did not erase the value of precise models for the tasks they fit.

Why GOMS belongs in HCI history

GOMS belongs in HCI history because it helped establish that the human side of a computer system can sometimes be modeled with enough precision to guide engineering. Goals, operators, methods, and selection rules gave analysts a structured description of expert work; the Keystroke-Level Model turned that structure into practical time predictions; and related models connected user intentions to command syntax.[1][2][4] Project Ernestine later showed that such predictions could influence real operational decisions.[5] The framework’s limits are equally instructive. GOMS does not replace usability testing, ethnography, or studies of novice learning. It excels when behavior is practiced, procedural, and measurable. Its historical contribution was to make that bounded promise explicit: when an interaction can be decomposed into repeatable human operations, interface performance can be analyzed before intuition hardens into design.

RESEARCH / PROVENANCE

Works Cited

5 SOURCES
  1. 01
  2. 02
  3. 03
  4. 04
  5. 05

CodeHistory is a living archive. Citations document the evidence used for this edition; later evidence may refine the account.

Contribute / Corrections

Improve the record.

Use this moderated submission form to suggest a correction, provide a source, challenge a priority claim or identify a missing contributor. Submissions are treated as research leads, not automatically published comments.

Submit a research lead

Please do not submit confidential material or claims you cannot support.