Further Explorations on the Use of Large Language Models for Thematic Analysis. Open-Ended Prompts, Better Terminologies and Thematic Maps

Paper · Source
Reading and SummarizationDomain Specialization in LLMs

There is a nascent area, where scholars are approaching thematic analysis (TA) using LLMs, following the six phases developed by BRAUN and CLARKE (2006). TA is a qualitative method of analysis where the researcher labels (codes) portions of data with relevant meaning and then organises these codes/labels into patterns (the themes). BRAUN and CLARKE stipulated that TA encompasses the following phases: 1. familiarisation with the data; 2. initial coding; 3. identification of themes; 4. revision of themes; 5. renaming and summarising of themes; and 6. write-up of the results.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Why do LLM recommenders underperform collaborative filtering despite their capabilities? How should retrieval systems handle complex multi-step reasoning? Does abstract user knowledge outperform concrete interaction history in personalization? What mechanisms preserve shared understanding in evolving conversations? Is language model reasoning authentic and what causes models to reason? What compositional reasoning failures limit large language models despite scale? Do reasoning traces faithfully reflect actual model reasoning? How effectively can language models perform reasoning, especially combined with symbolic methods? Do language models reason through causal mechanisms or semantic associations? Can models improve accuracy without degrading reasoning quality? What capability trade-offs arise from domain specialization through fine-tuning? How can humans maintain meaningful oversight as AI systems become increasingly autonomous and complex? Can local safety checks guarantee system-level behavioral safety? Where and how do personality traits reside in language models? How does harness optimization generalize across different model architectures and domains?