Documentation

How this works

Where the content comes from, what happens to each document, how a question is answered and what the portal checks before it shows you an answer.

Where the content comes from#

The portal reads one collection: research papers, their supplementary files and video material, loaded into the portal's knowledge index by the people who run it. Nothing else is read. What you can search, ask and browse is exactly what is in that collection, and the Library shows the current count.

The illustrated version of this page, How this works under Help, shows the live figures for the collection and a diagram of the flow described here.

What happens when a document is added#

Every document goes through the same steps before it can be found:

  1. Text and tables are extracted from the file page by page, so a passage can later be traced back to where it sits in the paper. Video material is transcribed.
  2. Enrichment agents read the extracted text and write a plain-language summary and key takeaways, assign topic and study-design labels, label individual passages, and record the relations between the entities the paper mentions (conditions, genes, medications, researchers and institutions) for the knowledge graph.
  3. The document, its passages and its labels are indexed for retrieval by meaning and by exact term.

The original file is never altered. The generated fields sit beside it and are shown on the document page as generated fields, never as the paper's own words.

How a question is answered#

Four things happen between the question and the answer:

  1. The question is routed. The portal reads the question and chooses the retrieval configuration that suits it: an identifier or a bare term is an exact lookup that lists the documents; a question about choosing or dosing a treatment is a clinical decision that always checks contraindications and monitoring; a broad question is an evidence review grounded on full text; a question about what is newest is answered newest first with the year stated; a question that names a table, a data sheet or a protocol document reads the supplementary files beside the papers; everything else runs on the default configuration. The choice is shown beside the answer as a chip, and you can change it and ask again.
  2. The index returns the passages. Before anything is retrieved, the portal resolves the things the question names against the collection: an antibody or antigen, a named consortium, registry or network, a trial acronym, a quoted title, a cohort you describe, and the medications and syndromes it knows. Where a name identifies a small enough set of papers, retrieval is restricted to those papers, the way asking a question of a single document is restricted to that document, so a question about one antibody cannot be answered with a neighbouring cohort's figures: that cohort's paper is never in front of the answer at all. A name that titles many papers is a topic rather than a name and restricts nothing, and a medication on its own never restricts retrieval, because a drug name titles a laboratory study and a clinical trial alike. Where a part of your question is not answered by the papers the names resolved to, the paper that does answer it joins them. When the collection does not hold a study the question names, the answer says so rather than answering from a paper that only cites it.
  3. A question that asks for a number is answered one paper at a time. A question that asks for a rate, a proportion, an age or a comparison is first broken into its clauses - "compare brivaracetam and perampanel" is two questions, "how old were the participants and how many were female" is two clauses about one study - and each clause is resolved to the one paper that answers it, using the names it uses and the medications and conditions it mentions. Each clause is then answered from that paper alone, the way a question about a single document is answered, and the answers are put together so that every sentence carries exactly one citation: no sentence draws on two papers, because no part of the answer was written with two papers in front of it. Where a clause has no paper, the answer says so for that clause and answers the rest. Questions that ask what the evidence is for something, rather than for a number, are still answered across papers. A question that names no medication is never split by medication: "which medications are contraindicated in this syndrome" is one question, and where the collection holds a consensus statement or guideline for the condition you name, that is the paper it is answered from.
  4. The answer is written only from those passages. Nothing is drawn from general knowledge or from the internet. Every sentence that states a finding carries a citation to the passage it came from (an item in a list takes the citation of the paragraph it belongs to), and opening the citation shows that passage in the paper. A sentence carries a marker only for a paper that shares a distinctive phrase of it, not merely its vocabulary, so a phrase every paper in the field uses lends no marker; where a sentence is left with none, the answer names that sentence under itself rather than leaving you to count markers.

How the answer is checked before you see it#

Before an answer is shown, the portal checks it against the cited text, sentence by sentence:

  • Figures. Every number, percentage, dose and range in a sentence is first located in the cited paper, and the sentence or table row that carries it there must share the claim's own quantity - its outcome, the noun the figure measures or the name the question asked about - about the same outcome, at the same follow-up, with the same responder threshold and the same denominator. The denominator is read from the figure's own sentence or table row, never from elsewhere in the paragraph: a number the paper writes as the count behind a share ("19 patients (28%)") agrees with a cohort size the answer pairs with it when the two make that share, and an analysis set the answer names beside a figure the paper pairs no size with must be the paper's own words. Where one sentence reports two arms, each figure belongs to the arm its own phrase names, so a placebo arm's rate is never served as the drug arm's. The outcome must match exactly wherever the paper itself is exact: where a paper reports both "seizure freedom" and "continuous seizure freedom", one does not stand for the other. A figure the cited passage does not carry is looked for in the full text of the retrieved papers: where one of them carries it beside the same claim, the sentence is cited to that paper instead; where the figure is there but cannot be tied to the claim as the answer stated it, the sentence is removed and that paper's own sentence on the outcome you asked about is quoted in its place. A figure found nowhere means the sentence is removed, and the answer says that it was. Removal is applied to the answer, not only recorded under it: a figure the note names has left the page, and a figure that still stands somewhere in the answer, verified where it stands, is not named as removed. Where the check empties an answer altogether, the paper the removed figure was found in is read before anything is declined, and the sentence that carries the figure there - from its own results, not its introduction, discussion or tables - is quoted and cited in place of the refusal.
  • Populations. A figure is bound to the group the paper reports it for. Where the passage a figure was found in names a group of its own, the group your question asked about must be that group or narrower: a rate the paper reports for "patients with psychiatric comorbidity" is not the rate for "patients who switched from levetiracetam to brivaracetam", and a sentence that points back ("of these patients") is checked against the group the sentence before it named. When the question names a cohort, trial or study, the papers that cohort names are the only papers retrieval reads, so no sentence can carry another cohort's figure, and a figure the cited paper only quotes from other studies is removed rather than annotated. Where the question names no cohort, such a figure is kept but marked as second-hand, with the paper's own finding beside it - and the marked sentence never leads the answer. What counts as second-hand is judged by the words, not the section: a figure a paper states in its own voice ("our cohort", "this trial", "we found"), reports in its own abstract, or prints beside the group it counted ("physicians: n = 19, 100%"), is that paper's finding wherever the extraction placed it, and a paper with no results section of its own - a review, a consensus statement - is judged on those words alone. A finding the answer credits by name to authors who did not write the paper cited beside it ("Rajna and Veres showed ...") is that paper's account of earlier work, and is named as such under the answer whether or not it carries a figure.
  • Named studies. A sentence cited to the wrong paper is replaced by the named paper's own sentence only when that sentence carries the same figure at the same time point, quoted verbatim and cited; otherwise the sentence is removed, and a named paper the answer never cited is read directly before anything is declined. A denominator the answer pairs with a figure is checked as part of the figure: a pairing the paper contradicts is removed and said so, never rewritten, and a denominator is only ever added from the figure's own bracket or table cell. When the papers that answer one question describe different populations, each sentence says which paper it comes from; a protocol's planned recruitment is named as such beside the results paper's enrolment. A study the answer names that this collection holds no paper for, and that no cited paper mentions, has nothing behind it: that sentence is removed and the answer says so. Reference lists are cut out of every paper before the check reads it, so a title in a bibliography can never stand in for a finding.
  • Years and safety verbs. A year must come from a cited resource. A medication the answer calls contraindicated must be called that, by name, in a cited passage: the verb is read with the medication nearest it, so a passage calling a different drug contraindicated is not support, and a passage that only calls the drug "not recommended", or says it may aggravate seizures, does not carry the stronger word. Where the sources say something weaker, the answer says so and quotes what they do say. The same holds for "should be avoided", a boxed warning and "first-line", and a medication the cited sources flag is never dropped silently.
  • What rested on a removed sentence goes with it. When a sentence is removed, the conclusion drawn from it goes too, and so does the opening answer when nothing else left in the answer stands behind it. An answer left with nothing but the notes the check wrote is declined, with the closest matches, rather than shown.

While the answer is still streaming, its text is shown as unchecked (muted, with a "still streaming, the check follows" mark), its first complete sentence is checked against the papers retrieval found and, when it passes, the paper that carries it is named under the answer; the checked answer then replaces the streamed text. A follow-up in the same conversation carries the earlier answers' cited papers with it: a question about "that study" is answered from those papers, with their own paragraphs and tables in front of the generator, and a request to put the earlier answers in a table keeps every row: each cell is checked under the column heading above it, so a figure filed under the wrong outcome is caught, and any cell the check could not verify - a figure it could not tie to that row's source, or an analysis set name where the column asked for a figure - is marked "not verified" rather than the row dropped. Chat with a document runs the same check against that document's own text and shows the same badge.

These checks are plain text comparisons against the extracted text of the papers, with no language model in the loop, so the check cannot invent support. The confidence label under the answer is led by that check: an unverified figure, year or contraindication marks it low, removed sentences cap it at moderate, and high is earned only when every figure was found. The platform's own quality scoring of how well the answer addresses the question, how firmly it is grounded and how relevant the retrieved passages were can lower the label but never raise it, and is shown as the platform's self-assessment. The check decides: a fluent answer whose figures the cited papers do not carry is not shown as high confidence.

What you can do with it#

  • Search finds documents fast, with a short cited answer over them or the results alone.
  • Ask is the full conversation: a grounded, cited answer, follow-ups that keep the context, saved sessions and deep research for broad questions.
  • Library and the reader browse the whole collection and open any paper at the cited passage.
  • Chat with a document asks questions of one paper alone; its answers are checked against that document's own text and badged the same way.
  • Investigations gather evidence around a research question over time and synthesise it.
  • Generate writes a briefing, comparison, timeline or set of questions and answers from the collection, with references.
  • Assessment builds a knowledge check on any area of the collection.
  • The knowledge map shows the conditions, genes, medications, researchers and institutions in the collection and how they connect.
  • Watches re-run a search or a question daily and flag it when the collection has something new.
  • Exports take an answer trail, an investigation or a generated artefact out as a Word document, and a briefing as a print-ready copy for saving as a PDF (portable document format) file.

What it deliberately does not do#

  • It never answers without a source. An answer with nothing to cite is not shown.
  • It says plainly when the collection does not hold something, and shows the closest passages it found, rather than filling the gap.
  • It does not browse the internet. Every answer comes from the collection alone.
  • It does not change the papers. Extraction and enrichment sit beside the original, which stays exactly as published.

Under the hood#

For technical readers. The knowledge index, retrieval, answer generation, citations, the answer quality signal, the enrichment agents and the entity relations graph are provided by Progress Agentic RAG (retrieval-augmented generation), the knowledge platform the portal runs on. The portal adds the intent routing, the verification layer described above, and the reading tools: the reader, document chat, investigations, generation, assessment, watches and exports. The platform sits behind one retrieval interface in the portal, and the credentials for it never reach the browser.