Explanatory Debugging: Supporting End-User Debugging of Machine-Learned Programs

City Research Online

City, University of London Institutional Repository

Citation: Kulesza, T., Stumpf, S., Burnett, M., Wong, W., Riche, Y., Moore, T., Oberst, I., Shinsel, A. and McIntosh, K. (2010). Explanatory debugging: Supporting end-user debugging of machine-learned programs. Proceedings - 2010 IEEE Symposium on Visual Languages and Human-Centric Computing, VL/HCC 2010, pp. 41-48. doi: 10.1109/VLHCC.2010.15

This is the unspecified version of the paper.

This version of the publication may differ from the final published version.

Permanent repository link: https://openaccess.city.ac.uk/id/eprint/211/

Link to published version: http://dx.doi.org/10.1109/VLHCC.2010.15

Copyright: City Research Online aims to make research outputs of City, University of London available to a wider audience. Copyright and Moral Rights remain with the author(s) and/or copyright holders. URLs from City Research Online may be freely distributed and linked to.

Reuse: Copies of full items can be used for personal research or study, educational, or not-for-profit purposes without prior permission or charge. Provided that the authors, title and full bibliographic details are credited, a hyperlink and/or URL is given for the original metadata page and the content is not changed in any way.


Explanatory Debugging:

Supporting End-User Debugging of Machine-Learned Programs

Todd Kulesza1, Simone Stumpf2, Margaret Burnett1, Weng-Keen Wong1, Yann Riche3, Travis Moore1, Ian Oberst1, Amber Shinsel1, Kevin McIntosh1
1Oregon State University, 2City University London, 3Riche Design
kuleszto, burnett, wong, moortrav, obersti, shinsela, mcintoke@eecs.oregonstate.edu
Simone.Stumpf.1@city.ac.uk, yann@yannriche.net

Abstract

Many machine-learning algorithms learn rules of behavior from individual end users, such as task-oriented desktop organizers and handwriting recognizers. These rules form a “program” that tells the computer what to do when future inputs arrive. Little research has explored how an end user can debug these programs when they make mistakes. We present our progress toward enabling end users to debug these learned programs via a Natural Programming methodology. We began with a formative study exploring how users reason about and correct a text-classification program. From the results, we derived and prototyped a concept based on “explanatory debugging”, then empirically evaluated it. Our results contribute methods for exposing a learned program’s logic to end users and for eliciting user corrections to improve the program’s predictions.

1. Introduction

Machine learning techniques are increasingly used in software adapted to end users’ own data, such as SPAM filters, recommender systems, and predictive text tools. These applications generate rules of behavior that are statistically derived via a particular user’s idiosyncratic patterns of behavior. We refer to these generated rules as machine-learned programs. While such programs can become fairly accurate, due to their statistical nature, they also remain fallible.

Who can fix a mistake made by a machine-learned program? The machine learning specialist who wrote the generator algorithm cannot fix every generated program for each individual user. Only one person is in a position to judge the correctness of the generated program: the very end user from whom the program has been learned.

End users, however, are given little power to correct a learned program’s errors. For example, SPAM filters confine user corrections to implicit approval and explicit disapproval. The user is permitted to scold the algorithm when it is wrong, but cannot tell the system why it was wrong.

This situation partly exists because enabling end users to debug machine-learned programs is hard. Learned programs use complex logic and, as generated programs, have no “source code” to directly represent this logic. Nevertheless, end users are capable of providing descriptive corrections beyond the binary scoldings commonly available today.

We present and evaluate a new Explanatory Debugging approach to harness this capability. Our approach supports debugging of learned programs by an iterative exchange of explanations between the program and the end user: the program explains how it arrived at its decisions, and the user explains where, in that decision-making process, it went wrong. We call this “explanatory” because it supports debugging via a give and take of explanations relating to existing or new machine learning features based on the user’s natural descriptions of concepts. (Features are elements used by machine learning reasoning, e.g., words, punctuation, etc.)

1.1. Domain: Coding in qualitative research

Our domain is coding—labeling segments of a transcript with codes for analysis in qualitative research—a common task for social scientists and HCI researchers. Such codes are developed based on the study’s research questions, so a code set is rarely reused in its entirety. Research involving coding of subjects’ verbalizations is labor-intensive, requiring hours of painstaking work. If a computer could “learn” from early examples how to code the remainder of an experiment’s transcripts, the time saved could be enormous. We refer to this possibility as auto-coding.

This domain is ideal for considering end-user debugging of machine-learned programs for three reasons. First, auto-coding is representative of a popular domain (text classification) that figures heavily in machine learning applications, e.g. SPAM filtering and predictive text technology. Second, debugging the program’s coding is needed because most studies have only a few subjects, resulting in too little data to reliably train the program. Third, if users can teach the program how to code well, the timesaving will be significant. If, however, they spend too long fixing the machine, the effort might exceed the time it takes to code everything by hand. Thus, the exchange between the program and the user must facilitate an accurate mental model of the program’s logic and must enable the user to explain how the coding should be done, so the learned program can benefit from these corrections.

2. Related work

There are systems that try to auto-code, e.g. the TagHelper system. While TagHelper can be highly accurate with lots of training data, obtaining a large set of coded examples is both expensive (because manual coding is time-consuming), and unrealistic (because data sets in qualitative analysis are usually small).

For users to debug the learned program’s logic, they must be able to see it. Explanations of learned programs’ logic have taken a variety of forms, such as relating user actions and the resulting predictions, detailing why a program made a particular prediction, or explaining how an outcome resulted from user actions. Much of the work in explaining probabilistic machine learning algorithms has focused on the naïve Bayes classifier and, more generally, on linear additive classifiers because explanations of these systems are relatively straightforward. More sophisticated but computationally expensive explanations exist for general Bayesian networks. However, these explanations are limited to account for the learned program’s behavior and do not extend to accepting user corrections to adapt future behavior.

Debugging involves two-way communication; once the program explains its logic, there needs to be a way for the user to adjust it. Some research has begun to shed light on supporting end users in fixing simple learned programs. Other systems explore building a program from the ground up by allowing users to specify the features it should employ. Research has also aimed at supporting experienced users in debugging more complex ensemble, sequential and non-sequential classifiers. None of this work has explored how to successfully engage end users in a two-way exchange in which they can introduce new machine learning features to fix complex learned logic.

3. Study #1: Explanations in debugging

Following the Natural Programming methodology, we began with a formative study (Study #1). Natural Programming is a user-centered methodology for designing programming languages and systems. It investigates users’ existing mental models (descriptions of existing concepts and processes) for a given task, and avoids influencing how participants think they are expected to do said task. The new system is designed to fit the users’ existing mental models.

Using this methodology, we investigated:

RQ1: Natural Explanations: How do end users “naturally” describe how to fix machine-learned programs?

RQ2: Existing Mental Models: How do end users think machine-learned programs make decisions?

RQ3: Mental Model Mutability: Can new information change end users’ existing mental models of machine-learned programs?

3.1. Participants, procedure, and tasks

Nine Psychology and HCI students (five female, four male) participated in our study; none had any experience with machine learning. Five participants had coded transcripts before, and all were familiar with Excel (required to understand the transcripts’ content).

The pre-task introduction involved practicing coding to become familiar with the technique and our codes, and completing a background demographic questionnaire. For the main task, we asked participants to help improve the accuracy of a system by judging the correctness of each code, fix the code when necessary, and to explain their reasoning.

We gave participants coded transcripts on printouts which they could mark-up using pens, colored pencils, etc., as they preferred. We also asked participants to “think aloud” and recorded their verbalizations, prompting them if their remarks were unclear.

The first 30 minutes of the main task aimed at eliciting natural participant responses and probing their existing mental models (RQ1 and RQ2). Participants worked on coded transcripts without explanations, then answered how they believed the computer did make its decisions and what information it should use. The final 20 minutes aimed at determining how explanations might influence users’ existing mental models (RQ3). This involved a variant of the coded transcript with explanations, after which participants told us how they now believed the computer made its decisions.

3.2. Materials

The transcripts came from an unrelated study about debugging spreadsheet errors. Although we told our participants that a computer had coded these transcripts, they had been hand-coded by a researcher using four codes: Seeking Information, Information Gained, Information Lost, and None. To elicit participant corrections, we introduced errors by randomly changing 30% of the expert’s codes.

We used paper printouts instead of a software prototype to elicit participant corrections in any form participants deemed appropriate, thus avoiding a tool that would restrict their range of expression.

The explanations drew on relationships within segments twice as frequently as relationships between segments, but distributed other feature types and relationships evenly.

3.3. Analysis methodology

Four researchers established an initial code set for analysis of the marked-up printouts and study transcripts, extending a code set used to research simpler machine learning approaches. Two researchers iteratively coded small sections of a transcript, adjusting the code set to clarify application. Inter-coder reliability between the two researchers on the final code set (applied to a different, complete transcript) was calculated by the Jaccard index as 81%. Given this acceptable level of code robustness, the two researchers coded the transcripts and questionnaire data.

3.4. Results: How should the program reason?

We first consider how participants explained how a learned program should reason (RQ1). As participants worked on coded transcripts without explanations, they mainly discussed single or multiple words, punctuation, and entire segments of text. These information types hold three implications for enabling end-user debugging of learned programs.

The learned programs need to reason about entire segments as participants' emphasis suggests. Additionally, removing punctuation carelessly cannot be done when processing data. Finally, even though algorithms tend to deal with individual words, participants discussed word combinations frequently.

3.5. Results: How did the program reason?

In this section, we consider how participants thought the computer did reason (RQ2), emphasizing how the program’s explanations were able to refine participant’s mental models about its logic (RQ3).

Before explanations were provided, participants thought the computer made decisions based on the presence of single keywords. However, after working with the explanations, most participants’ mental models included more complex types of reasoning, such as sequential relationships and the probabilistic nature of the program.

4. An Explanatory Debugging approach

As per the Natural Programming methodology, we used the results from Study #1 to design our Explanatory Debugging approach. Recall that the elements of Explanatory Debugging are an interactive give and take of explanations relating to existing or new machine learning features based on the user’s natural descriptions of concepts. Our AutoCoder prototype instantiates this approach and supports all of the results from Study #1. The basic coding and reasoning functionalities, which provide the context for Explanatory Debugging, are as follows. AutoCoder allows users to code segmented text transcripts with the same predefined codes as in Study #1.

To explain the logic behind a user’s code assignment, the user can highlight single and consecutive words, plus punctuation. These explanations can be complex, introducing features the learned program did not use before. Participants expressed fixes in a variety of forms, including single words, word combinations, punctuation, segments, and relationships.

5. Study #2: How well did Explanatory Debugging work?

In order to investigate how Explanatory Debugging supports end users fixing machine-learned programs, we conducted an empirical study exploring the following research questions: RQ4: Effectiveness: Which kinds of information (logic, runtime, or both) enabled end users to most effectively debug the learned program? RQ5: User Attitudes: How did users react to the availability of logic and runtime debugging information?

We recruited participants from the local student population and nearby residents. A 30-minute hands-on tutorial taught participants the concept of coding, the codes, and the prototype’s functionalities. After each part, participants answered free-form questions about how they believed the program made its decisions, plus Likert questions regarding the usefulness of each widget and participants’ perceived accuracy of the program.

6. Conclusion

This paper presented a new Explanatory Debugging approach for debugging machine-learned programs. Explanatory Debugging supports exchanges about logic (to support debugging’s code inspection aspects) and about outputs (to support debugging’s testing aspects). Our prototype let users see why the computer produced the outputs it did and explain their corrections. Most important, when using the Explanatory Debugging runtime-only variant, participants improved their programs significantly more than the current state of the art. How to guide users toward the most helpful corrections remains an open question, but this paper illustrates that a substantive exchange between an end user and their learned program is viable for users and can lead to more accurate machine-learned programs.