Skip to content
DocAnalyticaRequest access

Document intelligence

Structured, verified, AI-ready data from complex documents.

DocAnalytica transforms complex documents into structured, verified, and AI-ready data across scientific, legal, and other knowledge-intensive domains.

PilotScientific literature, nutrition pilot: 194,755 abstracts screened → 16,309 claims. Legal documents are in development.

claims · 1 row · nutrition pilotunedited
subject
VITAMIN_D
object
BONE
direction
NO_EFFECT
design
RCT
population
UNCLEAR← see note 1
species
HUMAN
sample_size
1294
dose
10 µg/d
gene
null
variant
null
statement
[withheld] ← see note 2
A

Evidence grade

source · Jennings A, Cashman KD, Gillings R, et al. A Mediterranean-like dietary pattern with vitamin D3 (10 µg/d) supplements reduced the rate of bone loss in older Europeans with osteoporosis at baseline: results of a 1-y randomized controlled trial. Am J Clin Nutr. 2018;108(3):633-640. doi:10.1093/ajcn/nqy122

  1. 1. Shown as produced, not cleaned up. The title says “older Europeans”, so a reviewer would correct population from UNCLEAR to OLDER_ADULTS. When a document doesn't state something clearly, it is recorded as unclear rather than guessed.
  2. 2. The original-text column is withheld: this paper's licence wasn't confirmed, and unknown means no.

What it does

From unstructured documents to standardized, machine-readable datasets.

Our AI-powered platform extracts facts, claims, evidence, and relationships from unstructured documents, converting them into standardized, machine-readable datasets.

  • TraceableEvery claim is traceable to its original source.
  • AuditableEvery quality assessment follows transparent, auditable criteria.
  • Licence-awareEvery export respects applicable licensing and access restrictions.

Built for researchers, legal professionals, enterprises, and AI developers, DocAnalytica turns fragmented information into reliable data for research, analytics, knowledge graphs, and AI model development.

Researchers

Pilot

Query the literature as data: subject, outcome, direction, study design, sample size and evidence grade, each row cited.

Legal professionals

In development

Traceable, licence-aware data from legal documents: provisions, obligations and the sources behind them.

Enterprises

In development

Versioned datasets and an API for analytics and research teams.

AI developers

In development

Source-linked records and reinforcement learning feedback for knowledge graphs, retrieval and model training.

How it works

From source document to reliable data.

Our process combines AI with rigorous quality controls, so that what reaches you is consistent, traceable and fit for use.

  1. 01

    Collect

    Documents are gathered from established, reputable sources, along with the details needed to cite them.

  2. 02

    Understand

    Each document is read for the findings it actually reports, including null and negative results.

  3. 03

    Check

    Records are checked for consistency before they are kept. Anything that doesn't hold up is left out.

  4. 04

    Assess

    Each finding receives a quality grade under consistent, documented criteria.

  5. 05

    In development

    Deliver

    Records become licence-aware datasets, ready for analysis, knowledge graphs and AI.

Reinforcement learning

In development

A platform that learns from every document it reads.

DocAnalytica is also a reinforcement learning platform. Models improve through feedback grounded in verifiable sources, so quality rises as each new domain is added, and the same feedback is available to teams training their own models.

  • Learns continuously

    Extraction quality improves from feedback rather than staying fixed at launch.

  • Grounded feedback

    Learning signals are tied to source documents, not to opinion.

  • Built for AI teams

    Source-linked, quality-graded data suited to training and evaluating models.

Why it matters

A document is prose. A question is a query.

“What does the evidence say about magnesium and sleep in older adults, and how good is it?” Ten thousand summaries can't answer that. They're a pile of text.

A claim is a row: this subject, in this population, moved this outcome, in this direction, by this design, at this sample size. Rows can be filtered, counted, compared and traced back.

Null and harmful results are recorded as faithfully as benefits. A supplement that did nothing is exactly the kind of finding a summary drops and a dataset should keep.

Proven in practice

Already at work on the scientific literature.

194,755
abstracts screened
16,309
claims extracted
10,344
distinct papers cited

Nutrition pilot, September 2026. Machine-extracted; not clinically reviewed.

Licensing

A dataset a buyer's lawyer can approve.

Facts and our quality grades are exportable. Original text is exportable only when the source's licence permits it. When the licence isn't known, the answer is no.

  • Unknown means no

    A source whose licence can't be confirmed is never treated as permissive.

  • Retractions excluded

    Claims from retracted papers are flagged and excluded by default. A retracted trial still reads like strong evidence, which is the danger.

  • Documented exports

    Each export comes with a record of its contents, sources and licences, and a checksum.

Questions

Plain answers.

Is this medical or legal advice?+

No. DocAnalytica describes what documents report. Claims are machine-extracted and no clinician or lawyer has reviewed them. Check the source before relying on any row.

How accurate is it?+

It depends on the field. In the nutrition pilot, subject, study design, grade, species and citation are the most reliable. Outcome is the weakest: a spot check found roughly 6 in 10 outcome labels correct. We publish these figures rather than hide them.

Can I get the original text of a finding?+

Only where the source's licence permits it (for example CC BY or CC0). Otherwise exports carry the structured fields and bibliographic data.

How does this relate to Nutrient-X?+

Nutrient-X's evidence base is built with DocAnalytica. DocAnalytica is the same capability offered as a product of its own.

When can I use it?+

Datasets and the API are in development. Request early access and tell us what you'd query; that shapes what we build first.

Tell us what you'd query.

Early access is for teams who need reliable, source-linked data from documents: research, legal, analytics and AI groups. We'll write back personally.

Or email hello@docanalytica.org