Researchers
PilotQuery the literature as data: subject, outcome, direction, study design, sample size and evidence grade, each row cited.
Document intelligence
DocAnalytica transforms complex documents into structured, verified, and AI-ready data across scientific, legal, and other knowledge-intensive domains.
PilotScientific literature, nutrition pilot: 194,755 abstracts screened → 16,309 claims. Legal documents are in development.
Evidence grade
source · Jennings A, Cashman KD, Gillings R, et al. A Mediterranean-like dietary pattern with vitamin D3 (10 µg/d) supplements reduced the rate of bone loss in older Europeans with osteoporosis at baseline: results of a 1-y randomized controlled trial. Am J Clin Nutr. 2018;108(3):633-640. doi:10.1093/ajcn/nqy122
What it does
Our AI-powered platform extracts facts, claims, evidence, and relationships from unstructured documents, converting them into standardized, machine-readable datasets.
Built for researchers, legal professionals, enterprises, and AI developers, DocAnalytica turns fragmented information into reliable data for research, analytics, knowledge graphs, and AI model development.
Query the literature as data: subject, outcome, direction, study design, sample size and evidence grade, each row cited.
Traceable, licence-aware data from legal documents: provisions, obligations and the sources behind them.
Versioned datasets and an API for analytics and research teams.
Source-linked records and reinforcement learning feedback for knowledge graphs, retrieval and model training.
How it works
Our process combines AI with rigorous quality controls, so that what reaches you is consistent, traceable and fit for use.
01
Documents are gathered from established, reputable sources, along with the details needed to cite them.
02
Each document is read for the findings it actually reports, including null and negative results.
03
Records are checked for consistency before they are kept. Anything that doesn't hold up is left out.
04
Each finding receives a quality grade under consistent, documented criteria.
05
In developmentRecords become licence-aware datasets, ready for analysis, knowledge graphs and AI.
Reinforcement learning
In developmentDocAnalytica is also a reinforcement learning platform. Models improve through feedback grounded in verifiable sources, so quality rises as each new domain is added, and the same feedback is available to teams training their own models.
Learns continuously
Extraction quality improves from feedback rather than staying fixed at launch.
Grounded feedback
Learning signals are tied to source documents, not to opinion.
Built for AI teams
Source-linked, quality-graded data suited to training and evaluating models.
Why it matters
“What does the evidence say about magnesium and sleep in older adults, and how good is it?” Ten thousand summaries can't answer that. They're a pile of text.
A claim is a row: this subject, in this population, moved this outcome, in this direction, by this design, at this sample size. Rows can be filtered, counted, compared and traced back.
Null and harmful results are recorded as faithfully as benefits. A supplement that did nothing is exactly the kind of finding a summary drops and a dataset should keep.
Proven in practice
Nutrition pilot, September 2026. Machine-extracted; not clinically reviewed.
Licensing
Facts and our quality grades are exportable. Original text is exportable only when the source's licence permits it. When the licence isn't known, the answer is no.
Unknown means no
A source whose licence can't be confirmed is never treated as permissive.
Retractions excluded
Claims from retracted papers are flagged and excluded by default. A retracted trial still reads like strong evidence, which is the danger.
Documented exports
Each export comes with a record of its contents, sources and licences, and a checksum.
Questions
No. DocAnalytica describes what documents report. Claims are machine-extracted and no clinician or lawyer has reviewed them. Check the source before relying on any row.
It depends on the field. In the nutrition pilot, subject, study design, grade, species and citation are the most reliable. Outcome is the weakest: a spot check found roughly 6 in 10 outcome labels correct. We publish these figures rather than hide them.
Only where the source's licence permits it (for example CC BY or CC0). Otherwise exports carry the structured fields and bibliographic data.
Nutrient-X's evidence base is built with DocAnalytica. DocAnalytica is the same capability offered as a product of its own.
Datasets and the API are in development. Request early access and tell us what you'd query; that shapes what we build first.
Early access is for teams who need reliable, source-linked data from documents: research, legal, analytics and AI groups. We'll write back personally.
Or email hello@docanalytica.org