Document analysis is the systematic use of existing texts — policies, minutes, reports, correspondence, case files, archives — as research evidence. Its founding insight is that documents are not windows onto reality but products of it: every record was written by someone, for someone, to do something, under conventions that shaped what could be said. The method therefore always works on two planes at once. Documents are read for content (what they state about the world) and interrogated as artefacts (why this record exists, what work it performed, what it systematically omits). Studies that skip the second plane treat the filing cabinet as a witness, when it was always a participant.
Documents as data: what they can and cannot attest
Documents excel as evidence of what was recorded and how: decisions as minuted, categories as officially defined, the vocabulary an institution required of itself at a given time. They are unbeatable for chronology, for positions taken in writing, and for the evolution of official framings across versions.
They are weak, alone, as evidence of practice: the gap between the procedure manual and the ward, between the minuted consensus and the argument that preceded it, is a standing finding of field research. The document attests that the organisation wrote this — a fact worth having — not that the organisation did this.
The classic appraisal criteria remain the working checklist: authenticity (is it what it purports to be?), credibility (produced with what accuracy and sincerity?), representativeness (typical of what class of records — and what never got filed?), and meaning (clear on its own terms and in its conventions?).
Methodological anchors
The most-cited procedural statement is Bowen’s Document analysis as a qualitative research method, covering the skim–read–interpret sequence, the combination of content and thematic techniques, and documents’ five research functions (context, questions, supplementary data, tracking change, verification).
For health-policy work, the READ approach — ready your materials, extract data, analyse, distil — set out in Dalglish, Khalid and McMahon’s Document analysis in health policy research: the READ approach, is a practical protocol that scales from a dozen policies to several hundred.
The document-as-artefact sensibility — records as situated products with authors, audiences, and functions — is the documentary-research tradition’s core teaching, and Scott’s four criteria above come from it (A Matter of Record, Polity Press).
When documents carry a study
- Retrospective questions: how a policy changed across a decade; what an organisation knew, and when, as recorded.
- Inaccessible settings: closed processes (court records, inquiry files) or past events with no interviewable witnesses.
- Official framing as the object: how “risk”, “merit”, or “safeguarding” is constructed in the record itself.
- Triangulation duty: alongside interviews or observation, documents anchor recollection to record — and expose the gaps that are themselves findings.
- Not sufficient alone for questions about practice, experience, or informal process; there the documents scope the study that other methods complete.
A defensible workflow
Constitute the corpus like a sample
Define the document universe (types, producers, period), the inclusion rule, and the search/acquisition path — then report what could not be obtained. Absences are data: the missing minutes and the unfiled reports have causes.
Log provenance per document
Author, date, audience, purpose, version, and route into your hands. This metadata sheet is the artefact-plane analysis in embryo, and it will settle disputes later about what a passage can attest.
Appraise before analysing
Run the four criteria explicitly, at least for load-bearing documents. Drafts, leaks, and retrospective summaries each have distinct evidential value — decided now, not mid-argument.
Extract and code on both planes
Content coding as the question requires (themes, categories, chronology), plus artefact coding: genre, stated purpose, silences, and intertextual references (which documents cite, echo, or overwrite which). Version comparison — what changed between drafts — is often where the study’s best evidence lives.
Synthesise with claims matched to warrant
Write conclusions that respect the two planes: “the record shows X was minuted” versus “interviewees report Y occurred” versus “the divergence between them suggests Z”. Document analysis is strongest when its claims advertise their evidential type.
Worked example: how “community consultation” hollowed out
A planning researcher analysed 15 years of one city’s consultation practice: 214 documents — statutory consultation reports, council minutes, internal guidance, and three versions of the consultation manual — constituted by rule from the council archive, with refusals and gaps logged (two contentious projects’ files were incomplete; noted, and pursued via FOI).
Content coding tracked the official story: consultation counts rose steadily. Artefact coding told the second story: across manual versions, the definition of consultation migrated from “opportunity to shape options” to “opportunity to comment on the preferred option”; report genres standardised around a template whose “objections addressed” section was, in later years, boilerplate reused verbatim across projects (caught by cross-document comparison); and minutes stopped recording dissenting submissions individually after a 2016 template change — a silence with a date.
Interviews with eight planners triangulated the record: the template change was remembered as an efficiency fix, its effects unintended. The study’s conclusion — consultation was proceduralised into auditability while losing deliberative content — rested on version diffs, genre shifts, and dated silences: findings only the documents could supply, qualified exactly as far as documents warrant.
Common mistakes
- Reading records as reality. The minute is evidence of minuting; practice claims need practice evidence.
- Convenience corpora. Analysing what was easy to download, then generalising to the institution’s documentation.
- Skipping provenance. Undated, unattributed extracts quoted as if origin did not condition meaning.
- Ignoring genre. A press release and an internal memo obey different truth conventions; coding them identically flattens the evidence.
- Missing the silences. What the record systematically omits is often the finding; unlogged absences cannot become one.
- Quote-mining. Extracting vivid lines without the document’s function — the artefact plane — reduces the method to decoration for a prior argument.
Limitations
The record is produced by the powerful side of most relationships: institutions document; clients, patients, and publics are documented. Corpora inherit that asymmetry, and analyses should name whose voice the archive structurally lacks.
Survival bias compounds it — what was kept, digitised, and indexed is a curated remainder, and the curation criteria are rarely recoverable.
And documents cannot be probed: no follow-up question, no clarification. Ambiguity that an interview would resolve in a sentence becomes, in documentary work, either a limitation stated or an inference defended — never silently resolved.
Where software helps
Document corpora are the native habitat of AI-assisted analysis: hundreds of PDFs coded against content and artefact codebooks, term migrations tracked across versions and years, and every extracted claim linked to its exact source passage. Evidano handles that pipeline — upload the archive, apply the codebook, compare across document types and periods — and its quote-level traceability is what lets documentary claims carry their provenance into the report.
Appraisal and inference stay with the researcher: no tool knows that the 2016 template change explains the silence, or which absences in the archive are meaningful. The software reads everything; deciding what the record can attest remains the craft.
Topics
- document analysis
- documentary research
- policy analysis
- archival research
- textual analysis
- qualitative document analysis
- secondary sources
Other methods in content and framework approaches
Written guides are linked directly; the rest have a reference entry in the methodology directory.
Published research using these methods
Studies and evaluations where this family of method was applied with Evidano — the work, not the claim.
- Published research2026

A Stanford-led study used AI to cross-check its qualitative coding
A published mixed-methods study uploaded deidentified transcripts into Evidano to cross-check themes. Every AI code was reviewed by the researcher.
9 interviews analyzed
Getting Down to Facts, Stanford SCALE Initiative
- Published research2026

AI-assisted qualitative analysis just showed up in published research
A peer-reviewed 2026 Education Sciences study used Evidano (previously AILYZE) to refine its qualitative analysis, with full researcher oversight.
22 courses' open-ended SET comments analyzed
Education Sciences
- Published research2026

AI-assisted qualitative coding, used in a peer-reviewed study
Peer-reviewed, transparent, human-validated AI-assisted qualitative analysis.
39,788 open-ended comments coded
Public Personnel Management (Sage)
Keep reading
- Research MethodsDiscourse Analysis: what language is doing, not just sayingDiscourse analysis treats talk and text as social action. Interpretative repertoires, the action orientation, how to do it, and the difference from content analysis.
- Research MethodsSummative Content Analysis: from word counts to meaningThe two-stage method: quantify manifest terms, then interpret their contextual use. When counting words is defensible qualitative research and when it is just counting.
- Research MethodsDirected Content Analysis: coding with theory up frontThe deductive branch of content analysis: building a codebook from existing theory, handling data the categories cannot hold, and reporting what the theory missed.
