Home > Publications Search > Publication Detail

Using Llm-Based Filtering To Develop Reliable Coding Schemes For Rare Debugging Strategies

Conference Paper/Proceedings
Using Llm-Based Filtering To Develop Reliable Coding Schemes For Rare Debugging Strategies
Publication Year:
2024
Publication Source:
Communications In Computer And Information Science
Volume:
2278 CCIS
Funding Type:
CAREER
Author(s):
Fan, Aysa Xuemo; Liu, Qianhui; Paquette, Luc; Pinto, Juan
Supporting Project(s):

Identifying and annotating student use of debugging strategies when solving computer programming problems can be a meaningful tool for studying and better understanding the development of debugging skills, which may lead to the design of effective pedagogical interventions. However, this process can be challenging when dealing with large datasets, especially when the strategies of interest are rare but important. This difficulty lies not only in the scale of the dataset but also in operationalizing these rare phenomena within the data. Operationalization requires annotators to first define how these rare phenomena manifest in the data and then obtain a sufficient number of positive examples to validate that this definition is reliable by accurately measuring Inter-Rater Reliability (IRR). This paper presents a method that leverages Large Language Models (LLMs) to efficiently exclude computer programming episodes that are unlikely to exhibit a specific debugging strategy. By using LLMs to filter out irrelevant programming episodes, this method focuses human annotation efforts on the most pertinent parts of the dataset, enabling experts to operationalize the coding scheme and reach IRR more efficiently. © The Author(s), under exclusive license to Springer Nature Switzerland AG 2024.