Paper deep dive
Chat Debugging: An Exploratory Study of Human-AI Collaboration to Debug Analog Circuits
John Hu, Andrew Ash
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 8/5/2026, 4:43:32 AM
Summary
This paper presents an exploratory study on 'Chat Debugging,' where undergraduate students use public-domain Large Language Models (LLMs) to troubleshoot malfunctioning analog circuits on breadboards and PCBs. The study analyzes chat logs from students under exam conditions to evaluate the effectiveness of LLMs in providing debugging suggestions, identifying strengths such as domain knowledge and hypothesis generation, while highlighting gaps like limitations in 2D/3D image reasoning and students' deficits in critical thinking.
Entities (10)
Relation Signals (7)
Chat Debugging â uses â Large Language Models
confidence 98% · troubleshooting malfunctioning analog circuits... through conversations with public-domain large language models (LLMs)
Chat Debugging â isstudiedat â Oklahoma State University
confidence 95% · This study took place in the laboratory sessions of a third-year undergraduate course... at Oklahoma State University
Chat Debugging â targets â Analog Circuits
confidence 95% · troubleshooting malfunctioning analog circuits on breadboards and printed circuit boards (PCB)
Chat Debugging â isconductedby â Undergraduates
confidence 92% · by undergraduates through conversations with public-domain large language models
Chat Debugging â ispartofcourse â ECEN 3314
confidence 90% · This study took place in the laboratory sessions of a third-year undergraduate course: ECEN 3314
Chat Debugging â evaluates â LLMs' limitations in 2D/3D image-based reasoning
confidence 85% · identified major gaps in LLM technologies... such as LLMs' limitations in 2D/3D image-based reasoning
Debugging by Design â iscomparedwith â Chat Debugging
confidence 80% · Table I lists a few current approaches to hardware debugging education. Debugging by Design (DbD)... This work
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:This research paper describes an exploratory study on the effectiveness of Chat Debugging: troubleshooting malfunctioning analog circuits on breadboards and printed circuit boards (PCB) by undergraduates through conversations with public-domain large language models (LLMs). Through thematic analysis of students' voluntarily shared chat logs when debugging pre-determined buggy circuits under exam and time pressure, we discovered multimodal usage patterns by students and considerable domain knowledge and sensible debugging suggestions offered by off-the-shelf LLMs. Meanwhile, we also identified major gaps in LLM technologies and students' skills during human-AI collaborative debugging, such as LLMs' limitations in 2D/3D image-based reasoning, unjustified tone of confidence, and students' deficits in fundamental concepts and critical thinking.
Tags
Links
- Source: https://arxiv.org/abs/2608.02955v1
- Canonical: https://arxiv.org/abs/2608.02955v1
Trouble viewing inline? Open PDF directly â
Full Text
46,732 characters extracted from source content.
Expand or collapse full text
Chat Debugging: An Exploratory Study of Human-AI Collaboration to Debug Analog Circuits John Hu and Andrew J. Ash School of Electrical and Computer Engineering Oklahoma State University Stillwater, OK 74078, USA Email:john.hu,andrew.ash@okstate.edu AbstractâThis research paper describes an exploratory study on the effectiveness of Chat Debugging: troubleshooting mal- functioning analog circuits on breadboards and printed circuit boards (PCB) by undergraduates through conversations with public-domain large language models (LLMs). Through thematic analysis of studentsâ voluntarily shared chat logs when debugging pre-determined buggy circuits under exam and time pressure, we discovered multimodal usage patterns by students and con- siderable domain knowledge and sensible debugging suggestions offered by off-the-shelf LLMs. Meanwhile, we also identified major gaps in LLM technologies and studentsâ skills during human-AI collaborative debugging, such as LLMsâ limitations in 2D/3D image-based reasoning, unjustified tone of confidence, and studentsâ deficits in fundamental concepts and critical thinking. Index Termsâlearning technology, problem-solving, large lan- guage models, microelectronics, hands-on learning I. INTRODUCTION Expediting post-silicon debugging is extremely important for a semiconductor companyâs bottom line. Despite the best planning and simulation coverage before chip fabrication,bugs can and do escape onto silicon. When that happens, unlike software bugs that can be fixed with a patch, chipmakers would have wasted five to seven million dollars on fabrication costs and time [1]. Post-silicon debugging is also challenging, as there are so many moving parts, and not all of them are fully understood. It is estimated that a typical semiconductor project spends 35% to 50% of its time on debugging [2]. Due to the uncertainty in the debugging time before a new product can be released, debugging has also gained the notorious nicknameof the Schedule Killer [2]. However, such an important skill is heavily undertaught in college. A recent survey among US and international univer- sities found that very few schools offer any chip debugging- related curriculum [3]. Part of the reason could be that debugging is hard to learn and harder to teach [4]. Cognitively, debugging requires one to hypothesize about possible root causes [5]. However, an empirical study found that humans are not good at making more than a few hypotheses [6]. Debugging also benefits from experience [7]. However, a novice has none. Affectively, debugging is challenging because many students view bugs as personal failures and start avoiding the subject [8]. This work was supported in part by the National Science Foundation Award No: EES-2321255 and a DaVinci Fellowship from the DaVinci Institute. TABLE I COMPARINGCHATDEBUGGING WITH OTHERAPPROACHES TO MICROELECTRONICSDEBUGGINGEDUCATION DbD[9], [10][11]This work CognitiveCommon bugs€ challengesNew bugspp€ AffectiveFear€ ChallengesFrustrationpp€ Anxietypp€ Table I lists a few current approaches to hardware debugging education. Debugging by Design (DbD) [9], [10] was proposed by the University of Pennsylvania, where K-12 students pur- posefully crafted buggy circuit objects for their peers to debug. The benefit was that it transformed bugs from âfailure arti- factsâ to objects to learn with, and students reported comfort, mischievousness, and fun after such debugging exercises. The limitation was that it did not address how to fix new bugs, and students receiving the buggy artifacts reported frustration, feeling that DbD bugs were intentional and too hard to detect. The second is a domain-specific microelectronics debugging education intervention [11]. Common circuit mistakes are explicitly taught through mini-lectures, and fear of debugging is also addressed with a debugging cheatsheet. However, the limitation is that the cheatsheet cannot help with new problems, and debugging new problems tends to take a long time, which fuels studentsâ anxiety and frustration. To advance the state-of-the-art in hardware debugging edu- cation, this paper explores if we can tap into Large Language Models (LLMs) to teach students debugging. Specifically, this paper aims to study the effectiveness of studentsâ freelancing, unstructured conversation with LLMs (Chat Debugging) for analog circuit debugging. The rationale is: (1) AI providesa âvirtual experienceâ from big data training that fills the gap of human novice. When students face new bugs, their direct ex- perience may be insufficient. Chat Debugging allows students to tap into other peopleâs experiences delivered through the disruptive technology of AI and LLMs. (2) AI can also provide âemotional supportâ through chatbot-style personalized and immediate feedback in any otherwise frustrating process. (3) Human agency and critical thinking are naturally emphasized as students navigate AI hallucinations and errors. As a first step, this paper reports on a one-year pilot study of arXiv:2608.02955v1 [cs.HC] 3 Aug 2026 Product Definition & Specification Architecture & System Design Integrated Circuit (IC) & Packaging Design Verification Post-Silicon Validation Step 1Step 2Step 3Step 4Step 5 tape-out Revisions (Pass 2, 3, ...) until bug-free IC fabrication ($$$, time) Fig. 1. New IC product development cycle [12] Chat Debugging to understand how effective todayâs off-the- shelf LLMs are in helping students debug fundamental analog circuits on breadboards and printed circuit boards. With the fast development of AI technologies, it is highly likely that domain-specific fine-tuning, adaptation, or alignment of LLMs may not be necessary for undergraduate-level engineering education, like introductory microelectronics. In addition to technology evaluation, we are also interested in studentsâ usage patterns characterization, such as how and in what modality todayâs undergraduates interact with AI. The rest of the paper is organized as follows: Section I provides necessary background information on debugging, debugging education, and LLM in engineering education. Section I lays out the theoretical foundation for human- AI collaboration for debugging. Sections IV and V list the research questions and methodology. Section VI presents the results. Section VII discusses the educational implications of this study. Finally, Section VIII concludes the paper. I. BACKGROUND A. Microelectronics Debugging Debugging, otherwise known as troubleshooting, is an in- dispensable step in modern integrated circuit (IC) design and production. Figure 1 shows the product development cycle [12] for a new IC product in the semiconductor industry. Between Step 4 and Step 5 exists a high-stakes milestone calledtape-out. Before tape-out, everything about the chipâs design can be altered like software through modifications of design files, such as hand-drawn schematics (analog) or Verilog codes (digital). During tape-out, all design files are electronically sent to a domestic or overseas foundry, such as Intel or Taiwan Semiconductor Manufacturing Company (TSMC), and the production of the prototype IC begins. The tape-out process is final and irreversible. After tape-out,if any design bugs or mistakes are discovered, either before the final silicon arrives or during post-silicon validation (Step 6), a design revision will be required. Depending on whether only metal-layer changes [13] or full-layer mask modifications are needed to fix the bug, the company will need to redo tape- out, which may incur partial or full (up to two months) time and financial (upward of $5 to $7 million depending on the process node [1]) costs again. Due to the high stakes involved, engineers are motivated to run thorough verifications of the IC design (Step 4) prior to tape-out. Debugging at thepre-siliconstage primarily relies on Electronic Design Automation (EDA) methods, such as ensuring simulation coverage, intelligent debugging based on log files and error messages, and logic equivalence check- ing across different abstraction levels. However, due to the growing complexity of todayâs ICs, long simulation time for full-system verification, and practical concerns in time-to- market, pre-silicon verifications cannot drag on indefinitely. An increasingly large number of bugs do manage to escape to silicon, according to decades of market studies [14], [15]. Hence, there is an inevitable second phase of debugging calledpost-silicondebugging (which is part of Step 5: Post- silicon validation in Figure 1). Here, the priority is to identify the root causes of all post-silicon bugs quickly and with high confidence to guide the âPass-2â redesign. A rush toward revision without fully understanding their root causes is obvi- ously bad because it could lead to another revision down the road. A delay in Pass 2 is equally bad as it just delays time- to-market and ultimately, time-to-revenue. However, unlike pre-silicon debugging, where EDA methods are abundant [16], post-silicon debugging suffer from limited observability, repeatability, and controllability of the fabricated ICs [17]. Orthogonal to pre-silicon and post-silicon debugging, which involves verifications of a new chip product, there is also board-leveldebugging. Here, the concern is not that a new IC has undiscovered bugs, but end users of mature and provenly correct chips somehow create faulty circuits on breadboards or printed-circuit boards (PCBs) in their applications [18]. To aid the automatic debugging of breadboards and PCBs, some researchers have framed it as a human-computer interaction (HCI) challenge [19]. They created various semi- or fully automatic visualization tools to help amateur users see where they have unintentionally created open, shorts, or misconnec- tions in their circuits, as well as offering additional bench instrumentation capabilities through add-on boards [20]. B. Debugging Education Despite the importance of debugging in IC development, debugging as a hardware skill remains undertaught in college. Arguably, the closest STEM field where debugging is highly valued in practice but undertaught in the curriculum is com- puter science, specifically in programming and software engi- neering. In Computer Science (CS) education, there has beena long history of studying the nature of debugging. For example, McCauleyet al.[21] systematically reviewed CS education literature and summarized its findings by why bugs occur, what types of bugs occur, the differences between experienced and novice debugging processes, and implications on how to teach computer program debugging in K-12 and university contexts [22]. Some of the consensus include that debugging skills do not necessarily follow subject matter understanding [23],and debugging as a skill should be explicitly taught [24]. Some of the more recent studies [25] include more rigid studies of intervention effectiveness, such as their impact on studentsâ debugging self-efficacy and debugging performance. Compared with programming debugging, hardware debug- ging education is much less explored and has only recently caught researchersâ attention. Romeoet al.[26] reviewed empirical studies and curricular interventions on non-CS de- bugging education following the PRISMA 2020 guidelines and found a need for rigorous scientific studies in the science and engineering education context. Among the very few rigorous studies, Fieldset al.[9] found that peer-to-peer debugging cultivated productive emotions, such as comfort, fun, and empathy, among high school students through interviews and qualitative observations. Duweet al.[27] reported the emer- gence of a developed debugging mindset after two semesters of sequential debugging education training among computer en- gineering major college students. Mehrabanet al.[28] framed microelectronics debugging education as the next Million- Dollar question for the semiconductor industry. Ashet al.[29] developed the first debugging performance instrument. Their actual intervention [11] remains a work in progress, though the overall debugging performance were moving in the right direction. C. LLMs in Engineering Education The fast-developing capabilities and accessibility of gener- ative artificial intelligence (GenAI) and large language model (LLM) technologies, such as ChatGPT, are revolutionizing all areas of education, including engineering education in univer- sity settings [30]. It is widely agreed, based on a systematic review of empirical studies [31], that integrating LLMs into teaching and learning environments has offered personalized learning experiences for students [32], automated assessment for instructors, and intelligent tutoring for diverse student needs. For example, instructors can build core course problems in chemical engineering through ChatGPT [33]. Chenet al. [34] also benchmarked different LLM models in the ability to complete homework assignments in circuit analysis in electrical engineering education and utilized LLMs to analyze what problems students struggle with the most. In light of the growing capabilities of LLMs, it is not surprising that some educators called for a complete benchmark of all undergradu- ate engineering curricula and their assignments against LLMâs capabilities [35] and rethink the learning objectives and value we are providing to students through a college education. Despite the attention on LLMs as a chatbot, assessment tool, and intelligent tutoring, there is relatively little attention to LLMs for debugging or debugging education. The selective few examples in this area include work by Maat al.[36], where they utilized LLMs to assimilate beginner students, who Understand Specifications Testing Hypothesis Formation Hypothesis Testing Repair Error Begin End Error? Y N Locate Error Human Pain Points Fig. 2. Katzâs [5] cognitive model for troubleshooting also have buggy code problems. By asking real students to play the role of teaching assistant (TA) and teach LLM âstudentsâ how to create comprehensive hypotheses and ultimately arrive at the right root cause, human studentsâ hypothesis generation skills were slightly improved. The closest example to our vision is ChatDBG [37], an AI-powered debugging assistant for C/C++/Python code debugging. However, such AI-based conversational debugging assistants as an educational effort has yet to be translated into engineering education contexts. I. THEORETICALFRAMEWORK To appreciate the enormous potential of human-AI collab- oration for analog circuit debugging, it is helpful to first dive into psychology and understand the pain points of humans troubleshooting alone. In this paper, the cognitive model of human troubleshooting serves as the theoretical foundation for our exploratory study. A. Cognitive Task Analysis of Debugging Cognitive Task Analysis (CTA) is a branch of applied psychology that uses a series of qualitative methods to yield information about the knowledge, thought processes, and goal structures that underlie experts at work [38]. Earlier work on CTA for general troubleshooting modeled debugging as a multi-step process: Katzet al.[5] modeled it as understanding the system, testing the system, locating errors, and repairing errors (shown in Figure 2). Gilmoreet al.[39] further elab- orated that locating errors usually involves repeated iterations of hypothesis formation and hypothesis testing. Johnson [40] modeled it as problem space construction, problem space reduction, hypothesis generation/testing, and solution gener- ation/testing (shown in Figure 3). Axtonet al.[41] later pro- posed a three-phase model. Schaafstalet al.[42] characterized debugging as four sub-tasks: formulate problem description, generate causes, test, and validate. Apain pointwithin Katzâs Construct Problem Space Reduce Problem Space Hypothesis generation and Testing Solution verification Begin End Human Pain Points Fig. 3. Johnsonâs [40] cognitive model for troubleshooting cognitive model ishypothesis generation. Alaboudiet al. [6] found that humans generally create no more than a few hypotheses per problem. They also found that when given a set of potential hypotheses, programmerâs debugging success rate and efficiency significantly increased. However, none of these early studies included experiences in their cognitive model, even though it has been widely recog- nized that experts draw heavily from their experience to speed up the troubleshooting process. To adequately capture the role of experience, Jonassenet al.[7] proposed a troubleshooting learning architecture consisting of three essential components: a multi-layered model of the system that includes topographic, function, strategic, and procedural representations, a simulator to test hypothesis, and a case library that stores relevant past experiences as advice for the learner. For experts, experienced- based debugging is the first method they use because it is also the most efficient [7]. For novices, however, thelack of expe- riencebecomes anotherpain point. Seeking a knowledgeable colleagueâs opinion can help, but such experts may not always be available. The thirdpain pointof human debugging alone is the risk ofworking memory overflow, as described by Schaafstalet al[42]. Schaafstalâs CTA is better understood using Johnsonâs cognitive model, shown in Figure 3. In Johnsonâs model, the second step is the reduction of the problem space [40]. The search for a likely root is more effective if large parts of the problem space can be discounted early on (âpruning of the search treeâ.) In order to be able to search selectively, it is essential that representation of the system be highly structured either with a functional or topological hierarchical model. However, inexperienced engineers may struggle to form such a conceptual representation and debug by trial and error. As the problem space grow, this random search may soon lead to humans losing track of where they were in the problem- solving process and why they were performing certain tasks, resulting in back up in the troubleshooting process [42]. B. Why could human-LLM collaboration be a game-changer? First, experience offers a shortcut to the debugging process (not shown in Figures 2 or 3). LLMs were trained on massive text corpora. If the text corpora include any hardware debug- ging experiences, LLMs can likely learn from them. During experience recall, humans may forget, but machines remember everything. Second, humans may struggle in hypothesizing generation and problem space reduction due to cognitive over- load. In contrast, generating capabilities, reasoning/planning, and short-term memory are all tasks that GenAI excels. Third, analog circuits lack a universal HDL. The NLP capability of LLMs allows humans to bypass formal languages and use human language to describe analog circuits without loss of information. Thus, human-LLM collaboration combines the strengths of both parties, charting out a new frontier for analog circuit debugging and debugging education (Table I). IV. RESEARCHQUESTIONS Motivated by the vision that human-AI collaborative debug- ging could overcome most, if not all, pain points of human debugging alone, this research seeks to explore how effective todayâs off-the-shelf LLMs are in helping students debug fundamental analog circuits on breadboards and PCBs. As foundation models continue to evolve, we hypothesize that even without domain-specific fine-tuning, many commer- cial or open-source LLMs may already come with sufficient analog domain knowledge to provide the guidance and as- sistance shown in Table I. If so, educators can encourage undergraduate students to use their preferred LLMs to assist in their debugging tasks. Specifically, this exploratory study has the following research questions (RQ): 1) What are studentsâ usage patterns of LLMs when de- bugging analog circuits? 2) What are some of the observed strengths of LLMs in guiding studentsâ debugging process? 3) What are some of the gaps in LLM technologies or stu- dentsâ skills during human-AI collaborative debugging? V. RESEARCHMETHODOLOGY Research Context:This study took place in the laboratory sessions of a third-year undergraduate course: ECEN 3314: Electronic Devices and Applications, using theSedra & Smith (Microelectronic Circuits) textbook, in the School of Electrical and Computer Engineering at Oklahoma State University (OSU), Stillwater. This course is a four-credit-hour (three- credit-hour lectures and one-credit-hour laboratory) mandatory course for all Bachelor of Science in Electrical Engineering (BSEE) and Computer Engineering (BSCpE) students. This study spans over Spring and Fall 2025. The average enrollment for the course is 60 and 40 students, respectively. Research Participants:All students enrolled in the class can choose to participate or not participate in the study. Thereis no penalty for students whether they choose to participate in the study. A graduate research assistant administrates and keeps the consent forms while the instructor is out of the room. The instructor does not know who has participated in the study TABLE I STRENGTHS ANDWEAKNESSES OFHUMAN ANDAIINANALOGCIRCUITDEBUGGING HumanAIHuman-AI collaboration Problem space€ObservationspCannot âreadâ hardware construction €NL descriptions€NLP capabilities.Domain understanding Problem spacepcognitive load€Reasoning reductionpstatus tracking€Planning.Loss of alignment pexperience recall€Pattern recognition Hypothesispcognitive fatigue€Hypothesis generation.Hallucination generation€Hands-onpCannot test hypotheses & testingmeasurementswithout experiments Repair testing€end-to-end testpCannot test hardware€Productivity Gain until the semester is over and all grades have been submitted. LLM conversation logs of students who did not participate in the study were removed from the Chatlog pool before this study. The research and participant recruitment protocol was approved by OSU IRB No: IRB-24-454. Data Sources:To evaluate the effectiveness of off-the- shelf LLMs in helping students debug analog circuits, we base our study on a lab final exam, which was a 30-minute timed hands-on exam that asked students to identify bugs on randomly given buggy circuits [29]. Students were also asked to demonstrate a fix to the problem to the proctoring instructor or TA in order to gain full credit. Accompanying the hands-on portion is a three-page written question and answer worksheet. Students have to answer three questions on paper: (1) What are the symptoms of the circuit, (2) What are the root causes, (3) What are some possible fixes. The worksheet will also have a line for the instructor to record how long it took a student to finish the whole problem. Students were asked prior to the exam whether they pre- ferred to use an LLM during the exam. All students who asked to use an LLM were granted permission. The only requirement is that they share their full Chatlog to the instructor rightafter the exam. Studentsâ Chat logs, together with their debugging exam answer sheets, form the data sources for this study. Data Analysis:This study primarily uses qualitative meth- ods [43] to study the Chatlogs to infer the usage patterns of students and the effectiveness of LLMs in helping students solve each buggy circuit problem. We adopted an inductive thematic analysis (TA) [44] to analyze two semesters of chat logs. Themes were identified at both the technology and psychological levels, capturing not only the factual correctness but also the underlying trust in AI suggestions. Given that we also have quantitative data, a comprehensive mixed-method approach would be our future work. VI. RESULTS Figure 4 shows the complete problem set used to evaluate human-AI collaboration in analog circuit debugging. The prob- lem set was conceived by the instructor in conjunction with multiple TAs over two years [29], and not all problems were offered to students in any given semester. Some problems (e.g., P4) evolved from a previous version (P3) due to a laboratory TABLE I CHATLOGDISTRIBUTION OVERPROBLEMS Problem SetSpring 2025Fall 2025 P1: Floating body on CS amplifier04 P2: Opamp IC upside-down12 P3: Improper bias: CE amplifier30 P4: Improper bias: CS amplifier02 P5: Flipped diode on voltage doubler10 P6: 1x-10x Probe setting04 Total512 equipment change. Other problems (e.g., P5) was dropped from Spring to Fall due to studentsâ tendency to rebuild instead of debug existing circuits. Finally, some problems (e.g., P2, P6) went through different PCB designs, though the principle of the bug remains unchanged. Table I showed the distribution of chat logs we received over two semesters. Since students voluntarily chose to use LLMs or not, and they were randomly assigned a problem, even though we provided approximately equal benches for each problem, the distribution of available chat logs for research for each problem was random. However, there was a general tendency of more students embracing LLMs for debugging from Spring to Fall, despite the class enrollment decreased (60 vs. 40) over the same period. A. RQ1: Studentsâ Usage Pattern Before the study, we expected students to use verbal descrip- tions to describe circuits and symptoms during the problem space construction phase (Table I). In practice, however, we noticed a non-ignorable usage pattern: Theme 1:Students used images to capture all context, including the physical artefact (circuits) and the assignment itself (exam questions) Figure 5 shows an example in Spring 2025 where Student A uploaded two images for ChatGPT-4o to debug. 2 out of 5 chat logs in Spring 2025 and 5 out of 12 chat logs in Fall 2025 followed this usage pattern. We suspect that the examâs time pressure and the convenience of photos both contributed to this usage pattern. The rest are either entirely verbal or include only images of the assignment, but not the circuit. P1 out V in V D V D R Bug: floating body (breadboard) out V in V D V D R Fix: connect NMOS Body to GND P2 Bug: IC pin 1 facing down (PCB) Fix: rotate IC or reconnect power/GND C+ V C- V C+ V C- V P3 1kHz 10mVpp DC offset: adjustable out V in V C V =5V C R 1kΩ 1N3904 Bug: DC offset too low Fix: DC offset appropriate out V in V D V D R 1 R 2 R Bug: (breadboard) 1 R >> 2 R Fix: pick a new to bias in the saturation region 1 M 1 M out V in V D V D R 1 R 2 R 1 M 1 R P4P5P6 Oscilloscope setting: Probe: 10x Probe setting: 1x Bug: 1x/10x setting mismatch Fix: Make both setting agree 2 D in V out V Bug: diode wrong polarity (breadboard) Fix: rotate 2 D 2 D 1 D Fig. 4. The Full Buggy Circuit Problem Set used over two semesters Student:uploaded an image ChatGPT 4o (4/23/2025) Is this a common emitter amplifier circuitStudent: ChatGPT: (Student took a photo of the debug exam question on paper) Yes, this is indeed ... Student: uploaded the following image: What is wrong with this ... from what you can seeStudent: listed 4 potential issues: 1. Collector resistor position 2. Transistor orientation 3. input signal setup 4. lack of biasing ChatGPT: provided actions: âTo verify or Fix:â 1. Verify collector goes to Vcc 2. Make sure emitter is GND 3. Add DC bias with AC input 4. Confirm E-B-C location Student: âWant me to sketch out a correct wiring diagram for your breadboard?â Yes, and where is my emitter located ChatGPT:(chain-of-thought reasoning) true root cause! TO-92 package flat side facing right emitter at bottom Correct connection If only students go back Fig. 5. Student Aâs chatlog showing image-based debugging (Theme 1) and LLM identifying P3âs true root cause among its top recommendations (Theme 2) B. RQ2: Strength in LLMs for Debugging Throughout the two semesters, we were pleasantly surprised by how well off-the-shelf LLMs can debug analog circuits. Without any domain-specific fine-tuning, ChatGPT and Gem- ini models were able to recommend the right actions for students. In some cases, chain-of-thought reasoning [45] was also demonstrated. Theme 2:LLMs were able to suggest true root causes among their zero-shot debugging recommendations. Figure 5 showed student Aâs chat log debugging P3. He first asked a clarification question: âIs this a common emitter amplifier circuit?â ChatGPT confirmed with a âYesâ, plus a long description of how the circuit works (skipped here). Then, the student continued the conversation, shown verbatim below: Student A:What is wrong with this BJT common-emitter amplifier from what you can see? ChatGPT:Good catch testing this on a breadboard! ... here are some possible issues of things to double-check: 1) Collector Resistor Position: The resistor (likely the collector resistorR C ) is not clearly connected be- tween the collector andV C . It looks like it might be going to the ground rail instead. 2) Transistor Orientation: The 1N3904âs leads may not be correctly oriented. The flat face of the 1N3904 should correspond to E (left), B (middle), C (right) when facing the flat side (check the datasheet). Mis- placing E/C will cause the transistor not to operate in active mode. 3) Input Signal Setup: It looks like the signal generator (black probe) may be connected to the emitter rather than the base. For a common emitter amplifier, the input should go to the base, and the emitter should be grounded. 4) Lack of Biasing Resistors: It seems youâre applying an AC signal directly to the base without DC biasing (no base resistor fromV C or voltage divider). That means the transistor might not be properly biased into its active region, making amplification unreliable. Thus, at zero-shot, the fourth hypothesis pointed to the true root cause. The follow-up between the student and LLM, which clearly demonstrated chain-of-thought reasoning: Student A:Yes, and where is my emitter located ChatGPT:Awesome â letâs get you sorted out. First, identifying the emitter in your setup: Youâre using a 1N3904 NPN transistor. Looking closely at your breadboard photo: I see a small black TO-92 package transistor near the center. Based on standard 1N3904 pinout, when youâre looking at the flat side facing you and the leads pointing downward: LeftâEmitter, MiddleâBase, RightâCollector. So, in your photo, the transistorâs flat face seems to be to the right. That would make: 1) Top leadâCollector (going to red clip with a resistor â likely correct). 2) Middle leadâBase (black clip going to signal generator). 3) Bottom leadâEmitter (going to ground). That means: Your emitter is the bottom-most lead in your setup, going to the black ground clip. That seems correctly grounded. Want me to generate a ....? Here, ChatGPT not only located the transistor of interest, but also identified the correct package (TO-92) and reasoned correctly about the relative E/B/C location. However, despite LLMâs confirmation that the connectivity seemed to be correct, Student A continued to suspect that there must be something wrong in the connections (based on his written responses) and therefore failed to debug the circuit within 30 minutes. C. RQ3: Gaps in LLMs or Student Skills Here, we are interested in the occasions where LLMs made mistakes that either misled the students or prolonged the debugging process. We identified three prominent themes: Theme 3.1:Almost all factual errors came after LLMs analyzed image inputs. We assumed such errors were due to LLMsâ limited ability to understand 3D structural and wiring relationships through breadboard or PCB images. ChatGPT 5.1 (12/01/2025) Student: uploaded the following image: Top of blue resistor is connected to C4, not C3. Black alligator clip is not on tied to â +â It is actually tied to â -â correctly through this yellow wire Takeaway: ChatGPT struggled in 3D spatial relationship Fig. 6. ChatGPT 5.1 struggled to recognized 3D spatial relationship Figure 6 showed Student Bâs debugging of P4. She also adopted Theme 1âs usage pattern and asked: Student B:Here is my circuit. There is a common source amplifier circuit and currently I am getting no output. What are my problems? ChatGPT:thought for 4m 33s. From the picture I can see several things that will keep this common-source amp from doing anything: 1) The âgroundâ clip is on the + rail. Down at the bottom-left, the black alligator clip is clearly bit- ing onto the rail that is marked â+â, not the blue ââ/groundâ rail.· · · ÂŹFix: Move that black alligator clip to the blue rail (ground)· · · 2) Drain resistorR D isnât clearly tied from VDD to the drain. The blue vertical resistor that is supposed to beR D looks like it goes from some random row to a row that does not line up with any MOSFET pin. ÂŹ Fix: Put one end ofR D directly in the +V rail· · · Here, ChatGPT 5.1 made several mistakes that could be traced back to its inability to reason or infer three-dimensional (3D) relationship between different wires based on a single image. First, the black alligator was not tied to the red â+â rail on the breadboard but the â-â rail correctly (Figure 6). ChatGPT 5.1 thought it was connected to â-â because the black alligator happen to sit righton top ofâ+â from the angle that the picture was taken. Second, the top of the blue resistor was not plugged into pointC3but ratherC4, effectively connecting to â+â through the horizontal, cyan wire on the left. ChatGPT 5.1 was again tricked into the wrong observation due to its inability to discern the âbendâ (Figure 6) from a top view. ChatGPT (unknown version) (12/05/2025) Student:uploaded the left & right photos at AIâs request Fig. 7. ChatGPT made visual mistakes while claiming â99%â confident Theme 3.2:The tone of confidence in LLMsâ recommenda- tions upon inspecting visual inputs is particularly concerning. Figure 7 shows Student Câs chat log debugging P6 with an unknown version of ChatGPT. Due to the nature of P6, student C had a PCB that was wired correctly and fully functional. The bug was in the oscilloscopeâs setting. The initial portion of the chat log proceeds smoothly, except for the fact that ChatGPT consistentlyasked forimage inputs. But as soon as ChatGPT received the left image in Figure 7, it started to generate overly confident statements: ChatGPT:Thanks for the photo â this helpsa LOT. Below is what I can seedirectly from the image, and what isalmost certainlycausing your âgainâ1â problem. § Your TL074 is wired to the WRONG PINS for the op-amp you think youâre using.· · ·Specifically, you are connecting to pins 5, 6, 7, not 1, 2, 3· · · If you want, send one more picture: If you show a straight-down shot when I can clearly see the pin numbering vs. your wires, I can markexactlywhere each connection should go. But Iâm 99% confident:wrong op-amp pinsâ wrong feedback loopâgainâ1. However, ChatGPT was wrong in its vision recognition. The board was connecting resistors and test leads to pins 1, 2, and 3, not pins 5, 6, and 7. Its tone of confidence is concerning, especially when it claimed, âI can mark exactly where each connection should goâ after a major visual recognition error. To compound on its error, after Student C uploaded the second photo in Figure 7, ChatGPT was more convinced: ChatGPT:Great âTHIS photo confirms exactly what the problem is· · ·What your PCB actually connects: 1) The large traces on the bottom go topads in the middle, not pins 4 or 11. 2) The supply rails (red clip = +15 V, black clip = â15 V) are attached to two pads that areNOT connected to the TL074 power pins. 3) Pins4 (VCCâ)and11 (VCC+)appear to beuncon- nectedon your board. All three statements above were false, as seen in Figure 7. Theme 3.3:Students lacking fundamental understanding of basic circuit concepts and offloaded critical thinking to AI. We arrived at this theme through multiple examples in studentsâ chat logs, debugging different problems. Some of them are shown below: Student D:What is the expected output of a common source amplifier (P1) Student E:should the amplified bjt signal be ac or dc (P3) Student F:so if my input probe is a 1x probe with a scaleing of 1V and my output probe is a 10x wiht a vertical scaling of 10 that is teh correct setup? (P6) Student D didnât know the correct behavior for P1. Student E was confused about AC and DC. Student F, despite being on the footstep of discovering the true root cause, confused oscilloscopeâs probe setting with the Volt/div knob and asked Gemini for help. On those occasions, students relied on AI for circuit fundamentals and lab skills, which is concerning. VII. DISCUSSIONS The observed studentsâ preference for image-based debug- ging presents both as a challenge and an opportunity. The challenge is that despite LLMsâ superb NLP capabilities, they are not inherently good at image recognition, especially when 3D spatial relationships are involved. The opportunity is that it demonstrates a need for domain-specific AI development, whether it involves agentic systems or 3D visual tool usage, to overcome LLMâs visual recognition limitations. For educators, the implications may be that we should encourage students to use AI as a conversational debugging guide, provided that we inform students of the imperfections in off-the-shelf LLMs in breadboard or PCB schematic recog- nition tasks. We should also caution students against LLMsâ potential unjustified confident claims. Furthermore, instructors should emphasize fundamental concepts, independent and crit- ical thinking, so that students take full control of their own debugging process. Finally, this study also has its own limitations. A sample size of 17, though appropriate for an exploratory study, remainsa small, self-selected group. More data collection may be needed to make the conclusion more general and robust. VIII. CONCLUSION This paper presents an exploratory study on the effectiveness of off-the-shelf LLMs in helping undergraduate students debug introductory analog circuits. In addition to text-based inputs, students also prefer using images to capture the buggy circuits. The LLMs were generally able to suggest true root causes among their zero-shot suggestions. However, overly confident claims are concerning, since LLMs were unable to determine whether they made any factual errors in 3D wiring recognition. Thus, the instructorâs guidance on these risks and emphasison circuit fundamentals and critical thinking are necessary. REFERENCES [1] A. Mutschler, âThe Problem With Post-Silicon Debug,â Feb. 2019. [Online]. Available: https://semiengineering.com/the-problem-with-pos t-silicon-debug/ [2] B. Bailey, âDebug: The Schedule Killer,â Jun. 2021. [Online]. Available: https://semiengineering.com/debug-the-schedule-killer/ [3] R. Sarmento, F. Pereira, and C. Felgueiras, âHardware Test subjects in academic education,â Jun. 2022, p. 1â5. [4] D. H. OâDell, âThe Debugging Mindset: Understanding thepsychology of learning strategies leads to effective problem-solvingskills.âQueue, vol. 15, no. 1, p. 71â90, Feb. 2017. [5] I. R. Katz and J. R. Anderson, âDebugging: An Analysis of Bug- Location Strategies,âHumanâComputer Interaction, vol. 3, no. 4, p. 351â399, Dec. 1987. [6] A. Alaboudi and T. D. LaToza, âUsing Hypotheses as a Debugging Aid,â in2020 IEEE Symp. Visual Languages Human-Centric Comput. (VL/HCC), Aug. 2020, p. 1â9. [7] D. H. Jonassen and W. Hung, âLearning to Troubleshoot: A New Theory-Based Design Architecture,âEduc. Psychology Review, vol. 18, no. 1, p. 77â114, Mar. 2006. [8] P. Nagvajara and B. Taskin, âDesign-for-Debug: A Vital Aspect in Edu- cation,â in2007 IEEE Int. Conf. Microelectronic Syst. Educ. (MSEâ07), Jun. 2007, p. 65â66. [9] D. A. Fields, Y. B. Kafai, L. Morales-Navarro, and J. T. Walker, âDe- bugging by design: A constructionist approach to high school studentsâ crafting and coding of electronic textiles as failure artefacts,âBritish J. Educ. Technol., vol. 52, no. 3, p. 1078â1092, 2021. [10] L. Morales-Navarro, D. A. Fields, and Y. B. Kafai, âGrowing Mindsets: Debugging by Design to Promote Studentsâ Growth Mindset Practices in Computer Science Class,âProc. 15th Int. Conf. Learning Sciences (ICLS), Jun. 2021. [11] A. Ash and J. Hu, âWIP: Exploring the value of a debuggingcheat sheet and mini lecture in improving undergraduate debugging skills and mindset,â in2025 IEEE Frontiers in Educ. (FIE) Conf., 2025, p. 1â5. [12] âChip Design and R&D,â Semiconductor Industry Association. [Online]. Available: https://w.semiconductors.org/policies/chip-design/ [13] K.-H. Chang, I. L. Markov, and V. Bertacco, âAutomatingpost-silicon debugging and repair,â in2007 IEEE/ACM Int. Conf. Computer-Aided Design (ICCAD), Nov. 2007, p. 91â98. [14] H. D. Foster, âTrends in functional verification: a 2014industry study,â inProc. 52nd Design Automation Conf. (DAC), Jun. 2015, p. 1â6. [15] H. Foster, âIC/ASIC Functional Verification Trend Report - 2024| Siemens Verification Academy,â Feb. 2025. [16] A. Nahir, A. Ziv, M. Abramovici, A. Camilleri, R. Galivanche, B. Bent- ley, H. Foster, A. Hu, V. Bertacco, and S. Kapoor, âBridging pre- silicon verification and post-silicon validation,â inProc. 47th Design Automation Conf. (DAC), Jun. 2010, p. 94â95. [17] S. Mitra, S. A. Seshia, and N. Nicolici, âPost-silicon validation opportu- nities, challenges and recent advances,â inProc. 47th Design Automation Conf. (DAC), Jun. 2010, p. 12â17. [18] J. Serritella, I. Williams, T. Green, and P. Rheinheimer, âBoard-level Troubleshooting,â Nov. 2018, Texas Instruments, Precision Labs. [19] E. Strasnick, M. Agrawala, and S. Follmer, âScanalog: Interactive De- sign and Debugging of Analog Circuits with Programmable Hardware,â inProc. 30th Annual ACM Symp. User Interface Softw. Technol. (UIST), Oct. 2017, p. 321â330. [20] T.-Y. Wu, B. Wang, J.-Y. Lee, H.-P. Shen, Y.-C. Wu, Y.-A.Chen, P.-S. Ku, M.-W. Hsu, Y.-C. Lin, and M. Y. Chen, âCircuitSense: Automatic Sensing of Physical Circuits and Generation of Virtual Circuits to Support Software Tools.â inProc. 30th Annual ACM Symp. User Interface Softw. Technol. (UIST), Oct. 2017, p. 311â319. [21] R. McCauley, S. Fitzgerald, G. Lewandowski, L. Murphy,B. Simon, L. Thomas, and C. Zander, âDebugging: a review of the literature from an educational perspective,âComputer Science Educ., vol. 18, no. 2, p. 67â92, Jun. 2008. [22] R. Chmiel and M. Loui, âAn integrated approach to instruction in debugging computer programs,â in33rd IEEE Frontiers in Educ. (FIE) Conf., vol. 3, Nov. 2003, p. S4Câ1. [23] C. M. Kessler and J. R. Anderson, âA model of novice debugging in LISP,â inWorkshop Empirical studies of programmers, 1986, p. 198â 212. [24] M. S. Carver and S. C. Risinger, âImproving childrenâs debugging skills,â inWorkshop Empirical studies of programmers, 1987, p. 147â 171. [25] T. Michaeli and R. Romeike, âImproving Debugging Skills in the Classroom: The Effects of Teaching a Systematic Debugging Process,â inProc. 14th Workshop Primary and Secondary Computing Education, Oct. 2019, p. 1â7. [26] C. L. Romeo and A. Olewnik, âTroubleshooting in Engineering Edu- cation: A Systematic Literature Review,â inProc. 2025 ASEE Annual Conf. & Expo., Jun. 2025. [27] H. Duwe, D. T. Rover, P. H. Jones, N. D. Fila, and M. Mina, âDefin- ing and Supporting a Debugging Mindset in Computer Engineering Courses,â in2022 IEEE Frontiers in Educ. (FIE) Conf., Oct. 2022, p. 1â9. [28] H. Mehraban and J. Hu, âBoard 293: How to Teach Debugging? The Next Million-Dollar Question in Microelectronics Education,â inProc. 2024 ASEE Annual Conf. & Expo., Portland, Oregon, Jun. 2024. [29] A. J. Ash, J. D. Cribbs, and J. Hu, âBoard 94: Work in Progress: Development of Lab-Based Assessment Tools to Gauge Undergraduatesâ Circuit Debugging Skills and Performance,â inProc. 2024 ASEE Annual Conf. & Expo., Portland, Oregon, Jun. 2024. [30] A. Johri, A. S. Katz, J. Qadir, and A. Hingle, âGenerative artificial intelligence and engineering education,âJ. Eng. Educ., vol. 112, no. 3, p. 572â577, Jul. 2023. [31] Y. Shi, K. Yu, Y. Dong, and F. Chen, âLarge language models in education: a systematic review of empirical applications,benefits, and challenges,âComputers and Education: Artificial Intelligence, vol. 10, p. 100529, Jun. 2026. [32] M. Menekse, âEnvisioning the future of learning and teaching engineer- ing in the artificial intelligence era: Opportunities and challenges,âJ. Eng. Educ., vol. 112, no. 3, p. 578â582, Jul. 2023. [33] M.-L. Tsai, C. W. Ong, and C.-L. Chen, âExploring the useof large language models (LLMs) in chemical engineering education:Building core course problem models with Chat-GPT,âEducation for Chemical Engineers, vol. 44, p. 71â95, Jul. 2023. [34] L. Chen, Z. Qin, Y. Guo, J. Rohde, and Y. Zhang, âBenchmarking Large Language Models on Homework Assessment in Circuit Analysis,âInt. J. Artificial Intelligence in Educ., vol. 35, no. 5, p. 3294â3355, Dec. 2025. [35] P. Jamieson, S. Bhunia, G. D. Ricco, B. A. Swanson, and B.V. Scoy, âLLM Prompting Methodology and Taxonomy to Benchmarkour Engineering Curriculums,â inProc. 2025 ASEE Annual Conf. & Expo., Jun. 2025. [36] Q. Ma, H. Shen, K. Koedinger, and S. T. Wu, âHow to Teach Program- ming in the AI Era? Using LLMs as a Teachable Agent for Debugging,â inArtificial Intelligence in Educ.Cham: Springer Nature Switzerland, 2024, p. 265â279. [37] K. H. Levin, N. van Kempen, E. D. Berger, and S. N. Freund,âChat- DBG: Augmenting Debugging with Large Language Models,âProc. ACM Softw. Eng., vol. 2, no. FSE, p. 1892â1913, Jun. 2025. [38] J. M. Schraagen, S. F. Chipman, and V. L. Shalin,Cognitive Task Analysis. Psychology Press, Jun. 2000. [39] D. J. Gilmore, âModels of debugging,âActa Psychologica, vol. 78, no. 1, p. 151â172, Dec. 1991. [40] S. D. Johnson, J. W. Flesher, and S.-P. Chung, âUnderstanding Trou- bleshooting Styles To Improve Training Methods,â inAmerican Voca- tional Association Convention, Dec. 1995. [41] T. Axton, D. Doverspike, S. Park, and G. Barrett, âA model of the information-processing and cognitive ability requirements for mechan- ical troubleshooting,âInt. J. Cognitive Ergonomics, vol. 1, no. 3, p. 245â266, 1997. [42] A. Schaafstal, J. M. Schraagen, and M. v. Berlo, âCognitive Task Analy- sis and Innovation of Training: The Case of Structured Troubleshooting,â Human Factors, vol. 42, no. 1, p. 75â75, Mar. 2000. [43] J. Saldana,The Coding Manual for Qualitative Researchers, 5th ed., Dec. 2024. [44] V. Braun and V. Clarke, âUsing thematic analysis in psychology,â Qualitative Research in Psychology, vol. 3, no. 2, p. 77â101, Jan. 2006. [45] J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. Chi, Q. V. Le, and D. Zhou, âChain-of-Thought Prompting Elicits Reasoning in Large Language Models,âAdv. Neural Inf. Process. Syst. (NeurIPS), vol. 35, p. 24 824â24 837, Dec. 2022.