Site Logo
How we trained our AI models

Purpose-built AI for qualitative research. Trained ethically, grounded in seminal scholarship.

Whether you do academic research, program evaluation, or market and product research, Evidano’s models are never trained on your uploads. They learn from a provenance-controlled mix of seminal methodological literature, open and permissioned corpora, and rights-checked repositories — matched to the methods you actually use.

Research domains
3
academia, evaluation, market & product
Methodologies supported
100+
each anchored to its seminal literature
Curated corpora & repositories
45+
provenance-tiered and rights-checked
Your data used for training
Never
uploads, prompts, and outputs excluded
The principle

Provenance first

Qualitative data is lived experience, testimony, customer voice, and community memory. A model trained indiscriminately on random internet text risks shallow coding, culturally insensitive interpretation, and unattributed reuse — the familiar problem of “garbage data in, garbage insights out.” Evidano fine tunes our AI models only on data whose licence or permission permits that use, and use restricted corpora only for evaluation, citation, or methodological grounding.

Permissible data only

Public-domain, CC0, CC BY, CC BY-SA, ODC-By, or explicitly permissioned data may be used for model fine-tuning, subject to its terms. Non-commercial, no-derivatives, research-only, clinical, Indigenous, and consent-gated data is treated as restricted. Culturally sensitive corpora are used in compliance with community norms, consent expectations, and non-extractive reuse.

Grounded in seminal scholarship

Every supported method traces back to the publications that established it — Braun and Clarke for reflexive thematic analysis, Glaser and Strauss for grounded theory, Noblit and Hare for meta-ethnography — extended through open-access papers that cite, apply, and refine those frameworks. The model learns each method’s boundaries, debates, and quality criteria, not a homogenized average of web text.

Purpose-built, not general-purpose

Corpora are matched to the tasks researchers actually perform: coding and memoing, theme development, CMO extraction, outcome harvesting, journey mapping, jobs-to-be-done analysis. The model can therefore analyze your dataset and produce themes, quotes, and codebook-grounded justifications you can audit.

Built for your use case

Three research domains, one standard of rigour

Evidano is fine-tuned for the methods, data genres, and evidence standards of each domain it serves. Switch domains to see the methodologies supported, how the training is matched to them, and exactly which data sources underpin that capability.

From reflexive thematic analysis to conversation analysis

Academic Research

Evidano supports the full landscape of qualitative traditions — thematic, grounded-theory, phenomenological, ethnographic, narrative, discourse, content, and framework approaches, plus the major evidence-synthesis designs. The model is trained to learn to respect each method’s boundaries: reflexive thematic analysis is not codebook analysis, semantic coding is not latent coding, and description is not interpretation.

Methodologies supported

Thematic Analysis
Reflexive Thematic Analysis (Braun and Clarke)Semantic Coding (surface-level meaning)Latent Coding (underlying meaning)
Evidence Synthesis
Systematic Literature ReviewQualitative Evidence SynthesisMeta-synthesisMeta-ethnographyNarrative SynthesisIntegrative ReviewScoping ReviewRapid Review
Grounded Theory
Classic Grounded Theory (Glaserian)Straussian Grounded TheoryConstructivist Grounded Theory (Charmaz)Dimensional Analysis (Schatzman)Situational Analysis (Clarke)Constant Comparative MethodGioia Method
Phenomenological Analysis
Interpretative Phenomenological Analysis (IPA)Descriptive Phenomenology (Husserlian)Hermeneutic Phenomenology (van Manen)Transcendental Phenomenology (Moustakas)Existential PhenomenologyEmpirical Phenomenological Psychology
Ethnographic & Observational
Classical EthnographyFocused or Rapid EthnographyInstitutional EthnographyNetnographyParticipant ObservationShadowing or Go-along ObservationAutoethnography
Narrative & Discourse Studies
Narrative InquiryNarrative Analysis (Labov, Riessman)Discourse AnalysisCritical Discourse AnalysisConversation AnalysisRhetorical or Genre Analysis
Content & Framework Approaches
Conventional Qualitative Content AnalysisDirected Content AnalysisSummative Content AnalysisFramework Analysis (Ritchie and Spencer)Document or Textual AnalysisTemplate Analysis

How the model is trained for academic work

  • Seminal methodological publications are paired with open-access papers that cite, apply, and refine them — so the model learns each method’s procedures, internal debates, quality criteria, and common misapplications, not just its definition.
  • Curated corpora of oral history, testimony, classroom discourse, academic speech, parliamentary debate, online deliberation, multilingual speech, and scholarly text ground the model in the genres academic projects actually contain.
  • Contradictory methodological positions are retained where they reflect genuine scholarly debate, so the model gives context-sensitive guidance rather than a single homogenized account of qualitative research.

Provenance-tiered corpora for academic discourse

Every academic corpus sits in one of four provenance tiers based on what its licence and ethical context actually permit. Only tier-appropriate use is made of each source.

Tier 1 · Open or public domain

Public-domain, CC0, CC BY, CC BY-SA, or ODC-By terms that permit training, with attribution and share-alike obligations honoured.

Tier 2 · Institutional research

Library, university, and consortium collections used within their research-access terms — for evaluation, citation, and methodological grounding, or under explicit licence.

Tier 3 · Research-only / NC-ND

Non-commercial or no-derivatives sources. Used as research reference and evaluation material only.

Tier 1 · Open or public domain10 sources

Tier 2 · Institutional research9 sources

Tier 3 · Research-only / NC-ND8 sources

Scholarly grounding

Every method traces back to its seminal literature

We began with the publications that established or substantially advanced each method, then expanded each lineage with open-access papers that cite the seminal work as a primary framework, apply its procedures, or refine it. Where a seminal text is restricted, the full text is never ingested — only lawful metadata, abstracts, and open-access operationalizations. Below are the anchor works behind the academic method families.

Evaluation and market-research method families — from realist evaluation and outcome harvesting to jobs-to-be-done and usability testing — are anchored the same way: foundational guidance documents and high-quality exemplars first, then open materials that apply and refine them.

How the corpus is built

A governed pipeline, from intake to retraining

Every source moves through the same documented stages and the guiding principle throughout is that AI should assist the researcher’s judgement, not obscure it.

01

Seminal-literature anchoring

For each method, curation starts with the publications that established or substantially advanced it, then expands through open-access papers that cite the seminal work as a primary framework, apply its procedures, clarify misunderstandings, or propose refinements. Where a seminal text is restricted, only lawful metadata, abstracts, and open-access papers that accurately describe and operationalize it. This aims to respect intellectual property and preserve scholarly attribution.

02

Rights, consent, and privacy gate

Licence, consent, and terms of use are confirmed before any data is included. If rights, consent, or privacy are unacceptable, the source is excluded regardless of its analytical value.

03

Quality scoring and weighting

Every candidate source is scored on a ten-dimension rubric before inclusion: relevance, methodological quality, evidence traceability, annotation quality, representativeness, diversity, ethical and legal clearance, recency, completeness, and risk. A small amount of expert-annotated data or seminal scholarship is weighted more heavily than a large amount of unreviewed public text.

04

Bias and representativeness review

Bias is audited at three levels. Data bias: what is over- or under-represented in the corpus, corrected through stratified sampling and targeted acquisition. Methodological bias: donor assumptions in theories of change, dominant voices in focus groups, platform effects in online data — flagged rather than treated as neutral evidence. Model bias: tested with subgroup performance, counterfactual examples, multilingual cases, and adversarial prompts. Demographics are not inferred from names or appearance; diversity is assessed from declared study context, setting, language, and methodological position.

05

Evaluation, expert review, and living governance

Models are benchmarked on accuracy, evidence grounding, methodological fidelity, uncertainty calibration, fairness across languages and populations, reproducibility, and auditability. Benchmark sets deliberately include contradictory evidence and cases where the correct answer is “insufficient evidence.” Human experts review outputs, every publication is linked to source metadata, and the corpus is treated as a living resource — periodically reviewed, expanded, and corrected as citation patterns change and communities identify omissions.

How sources are weighted

Training weight = quality score × task relevance × evidence traceability × diversity adjustment × freshness adjustment × risk penalty

Each candidate source is screened for ethical and legal clearance, and then scored across ten dimensions.

RelevanceMethodological qualityEvidence traceabilityAnnotation qualityRepresentativenessDiversityEthical & legal clearance (gate)RecencyCompletenessRisk
Bias, plagiarism, and quality

High-quality data

Strong data sources curated by libraries, universities, international organizations, and research consortia are used.

Ethics-first

A source that fails licence, consent, or privacy checks is excluded. Sensitive personal narratives are not trained on without a clear legal and ethical basis.

Provenance aware

Methods are learned from lawful metadata, abstracts, and open-access literature that cites the original frameworks, so attribution and provenance travel with the knowledge. Retracted papers and unclear-provenance sources are excluded.

Bias auditing

Bias is checked in the data (what is represented), in the method (how findings were produced), and in the model (what behaviour was learned), with subgroup, multilingual, and adversarial testing as standard.

Debate retained

Competing methodological positions are preserved where they reflect genuine scholarly disagreement. The models are evaluated on acknowledging uncertainty and avoiding claims of academic consensus where none exists.

Negative examples included

Weak theories of change, unsupported claims, quality assurance assessments, thin evidence, and misclassified themes are part of the corpus, so the model learns what poor analysis looks like, not just good analysis.

Security and reproducibility

Your research is never our training data

Secure AI is not only about encryption. Safely defend your analysis to a peer reviewer, an IRB, a funder, or a client.

What this means for you

  • Your confidential project data — transcripts, codebooks, prompts, and outputs — never trains the models.
  • The model is purpose-built for your use case, whether that is academic research, program evaluation, or market and product research — over 100 methodologies, each anchored to its seminal literature.
  • Fine-tuning training data comes from curated, provenance-tiered, rights-checked corpora (not indiscriminate web scraping).
  • Bias is audited at the data, method, and model levels.
  • Outputs are judged by transparent evidence trails, linked quotes, codebook consistency, and reproducibility.
Company
About
Newsletter

Product updates, research, and tips — straight to your inbox.

© Evidano, All Rights Reserved.