LLM-Assisted Thematic Analysis Protocol (LATA)
Phase 1 — Concept Discovery (v1.0)

1. Purpose

This phase aims to identify the latent thematic structure of a bibliographic corpus through an inductive, corpus-driven exploration of its textual content. Rather than assigning publications to predefined categories, the objective is to reconstruct the principal research phenomena that emerge from the literature and represent them as a set of coherent latent topics.

Concept Discovery constitutes the exploratory foundation of the protocol. At this stage, emphasis is placed on understanding the intellectual structure of the corpus while preserving sufficient thematic granularity for subsequent conceptual integration.

2. Input

The input consists of a bibliographic corpus containing one record per publication.

The thematic analysis should be based exclusively on the textual information available for each document:

Title
Abstract
Author keywords

Additional bibliometric metadata (e.g., publication year, journal, authors, affiliations or citation counts) may provide contextual information about the corpus but should not influence thematic identification.

3. Analytical Principles

This phase follows five methodological principles:

a) Corpus-driven discovery. Themes should emerge inductively from the corpus rather than from predefined disciplinary, political or theoretical frameworks.

b) Conceptual reasoning. Topics should represent shared research phenomena rather than recurring terminology. Conceptual similarity takes precedence over lexical similarity.

c) Progressive familiarization. The corpus should be understood holistically before proposing any thematic organization.

d) Appropriate granularity. Topics should be sufficiently specific to capture meaningful conceptual distinctions while avoiding unnecessary fragmentation.

e) Interpretative transparency. Conceptual uncertainty should be explicitly acknowledged rather than resolved through arbitrary classification decisions.

4. Operational Instructions

Conduct an exploratory thematic analysis as an experienced qualitative researcher seeking to understand the conceptual organization of the corpus.

The analysis should proceed through four sequential stages.

Stage 1 — Corpus Familiarization

Develop an overall understanding of the corpus before identifying thematic structures.

Identify recurring actors, institutions, organizations, geographical contexts, policy domains, social phenomena, research questions, and analytical perspectives.

Avoid drawing premature thematic conclusions.

Stage 2 — Conceptual Pattern Recognition

Examine the corpus to identify recurring conceptual patterns.

Focus on similarities in research objectives, phenomena under investigation, analytical approaches, and conceptual problems rather than on simple word frequency.

Treat documents as potentially multidimensional when they address more than one research phenomenon.

Stage 3 — Latent Topic Construction

Construct a preliminary inventory of latent topics.

Each topic should:

represent a coherent conceptual phenomenon;
encompass multiple related publications;
possess interpretable conceptual boundaries;
remain distinguishable from neighboring topics.

Maintain sufficient thematic granularity to support subsequent integration into higher-order conceptual structures.

Stage 4 — Internal Critical Review

Critically evaluate the proposed topic structure before completing the analysis.

Review the thematic framework to identify:

redundant or overlapping topics;
excessively broad or overly fragmented topics;
conceptually ambiguous publications;
underrepresented but meaningful research phenomena.

Where uncertainty exists, preserve it explicitly rather than forcing artificial thematic separation.

5. Expected Output

The output of this phase should include:

a concise conceptual overview of the corpus;
an inventory of latent topics;
a conceptual description of each topic;
representative concepts associated with each topic;
representative publications illustrating each topic;
observations regarding conceptual overlaps or thematic ambiguities.

The objective is not to produce a definitive classification but to generate a coherent conceptual representation suitable for higher-level thematic integration.

6. Quality Assurance Checks

6.1. Analytical calibration

Before accepting the latent topic structure, manually review a small but conceptually representative sample of document-level topic assignments. The sample should include clearly classified documents, conceptually ambiguous cases, and publications located near the boundaries between topics.

Evaluate whether the proposed latent topics accurately represent the principal research phenomena emerging from the corpus rather than recurring terminology alone. If systematic conceptual inconsistencies are identified, document representative examples, explain why the proposed topic assignment is inappropriate, refine the analytical instructions, and repeat the topic identification process. Proceed only when the topic structure demonstrates satisfactory conceptual consistency and interpretability.

6.2. Phase-specific checks

Before completing this phase, verify that:

the principal conceptual areas represented in the corpus have been identified;
topics are defined by conceptual meaning rather than lexical similarity;
thematic boundaries remain coherent and interpretable;
multidisciplinary publications have been appropriately considered;
conceptual uncertainty has been documented where appropriate;
the resulting latent topics provide an appropriate level of abstraction for subsequent macrocluster construction.

7. Transition to Phase 2

The validated inventory of latent topics constitutes the input for Phase 2 — Concept Integration.

The objective of the next phase is to examine conceptual relationships among latent topics and organize them into a coherent set of higher-order macroclusters while preserving the conceptual distinctions identified during the discovery stage.