Designing Against Deskilling: Metacognitive Feedback Reduces Cognitive Offloading to LLM Assistants

Paper · arXiv 2609.20143 · Published September 17, 2026
AI in Education

Cognitive offloading to AI can reduce opportunities to practice skills, creating risks of deskilling. However, it remains unclear how to prevent deskilling without restricting access to AI. Here, we design two interventions to reduce offloading decisions: (1) metacognitive feedback that makes the implications of offloading for users explicit, and (2) an effort-based reward that incentivizes less extensive LLM assistance. We test both in a preregistered online experiment (N= 704) with a 2×2 design and a no-AI control. The task was to practice fraction arithmetic with an LLM-based assistant that provided solutions only on explicit request, followed by an unaided test. Metacognitive feedback reduced answer offloading (OR = 0.47) and improved test performance (OR = 1.51). We found no evidence that the reward affected either outcome. Our results identify metacognitive feedback as a promising design choice to reduce cognitive offloading.

Introduction. LLM assistants offer a new way to support learning and skill development through natural language interaction. In educational contexts, such LLM assistants can explain unfamiliar concepts, provide feedback on reasoning, and generate individualized problems to practice, for instance by scaffolding programming practice [39] or supporting mathematics learning [55]. However, access to LLM support does not necessarily translate into better skills when the assistance is no longer available. There is growing concern that LLM assistance may even undermine skill development or erode existing skills, commonly discussed as deskilling [49]. For example, high-school students who practiced mathematics with unrestricted ChatGPT scored higher during practice yet lower on the subsequent exam without LLM assistance [5]. Similarly, experienced endoscopists detected fewer adenomas when performing colonoscopies without AI after months of AI-assisted practice [11], and LLM assistance during creative tasks lowered subsequent independent creativity [45].

Discussion / Conclusion. We set out to test whether interaction design can reduce cognitive offloading and thereby protect learners from deskilling. For this, we compared two interventions (i.e., metacognitive feedback and a reward) in a randomized 2 × 2 design with a no-AI control. Our results show that, under this design, there is no evidence that access to the AI assistant did impair subsequent unaided performance. Metacognitive feedback reduced answer offloading and improved unaided performance, while the reward affected neither outcome. Our exploratory analyses further clarify where the learning risk may arise. Participants who requested complete answers more frequently tended to perform worse on the subsequent test. Answer offloading also varied strongly across participants and was more common among those lower in need for cognition and perceived confidence. Together, these findings suggest that the relevant risk lies less in access to LLM assistance itself than in how much cognitive work learners choose to hand over. Metacognitive feedback appears to shift this decision by helping learners regulate their use of assistance while leaving the full range of LLM support available.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Does AI assistance promote real skill development or substitute for independent learning? What structural properties of attention create systematic model biases? Does warmth and empathy training systematically degrade model reliability? Why does polished presentation create unearned authority in AI outputs? How can humans maintain meaningful oversight as AI systems become increasingly autonomous and complex? Why is hallucination an inevitable limitation of current language models? When do multi-agent systems outperform single frontier models? What is the relationship between thinking tokens and reasoning accuracy? How should designers communicate what AI systems truly are and can do? What compositional reasoning failures limit large language models despite scale? What determines appropriate intervention timing and manner for AI agents? Can mechanistic interpretability reliably guide practical model design choices?