Études fondées sur les communautés Reddit

Automating inductive thematic analyses of health content using large language models: a proof-of-concept study using social media data.

Hairston J, Ranjan R, Lakamana S, Spadaro A, Bozkurt S, Perrone J, Sarker A

JAMIA Open . 2025;8 (5) :ooaf102

📅 01/10/2025 PMID : 40985037 DOI : 10.1093/jamiaopen/ooaf102

Résumé

OBJECTIVES: Large language models (LLMs) face challenges in inductive thematic analysis, a task requiring deep interpretive, domain-specific expertise. We evaluated the feasibility of using LLMs to replicate expert-driven thematic analysis of social media data.MATERIALS AND METHODS: Using 2 temporally nonintersecting Reddit datasets on xylazine ( = 286 and 686, for model optimization and validation, respectively) with 12 expert-derived themes, we evaluated 5 LLMs against expert coding. We modeled the task as a series of binary classifications, rather than a single, multilabel classification, employing zero-, single-, and few-shot prompting strategies and measuring performance via accuracy, precision, recall, and F score.RESULTS: On the validation set, GPT-4o with 2-shot prompting performed best (accuracy: 90.9%; F score: 0.71). For high-prevalence themes, model-derived thematic distributions closely mirrored expert classifications (eg, xylazine: 13.6% vs 17.8%; medications for opioid use disorders: 16.5% vs 17.8%).CONCLUSION: Our findings suggest that few-shot LLM-based approaches can automate thematic analyses, offering a scalable supplement for qualitative research.