Abstract
Development of Large Language Models (LLMs) offers new opportunities to analyse large qualitative data sets. This study proposes a framework that uses LLMs from five popular families (GPT, Mistral, Claude, DeepSeek and Llama) to automate data extraction process. The framework is tested on real estate listings from Luxembourg webscraped in 2025 (Jan) to demonstrate: (1) framework stability -ability to generate similar responses when provided with the same data set, but also of following the same pattern when faced with new questions, (2) framework calibration mechanism – by adding code sheets that clarify how to interpret specific situations, (3) mechanism of validating model behaviour over larger data set – by calculating correlations between partial responses generated by the framework and comparing them with the pattern generated on the test set.
| Original language | English |
|---|---|
| Publisher | SocArXiv |
| Number of pages | 14 |
| DOIs | |
| Publication status | Published - 12 Mar 2025 |
Bibliographical note
This article was submitted and deposit in arXiv : Open archive of the social sciencesKeywords
- GenAI
- LLMs
- data extraction
- qualitative attributes
- real estate
Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver