Biological input
Were the requested genes, species, disease, process, and unresolved terms represented correctly?
Biomedical GenAI Evaluation
This page uses a six-stage evaluation structure. It separates biological input, candidate retrieval, ranking, synthesis, citation support, and LLM-as-judge reliability instead of collapsing the system into one quality score.
Download evaluation JSON →01 · Stage-specific evaluation
Defined separate success criteria for selected genes, species, and biological context; evidence retrieval and ranking; answer generation; citation support; and LLM-as-judge reliability.
Were the requested genes, species, disease, process, and unresolved terms represented correctly?
Did reference publications enter the graph-derived candidate pool?
Were relevant candidates placed early enough to be inspected?
Did synthesis stay within the selected title-and-abstract evidence?
Does each cited PMID support the specific claim to which it is attached?
Does the automated judge follow its rubric, preserve evidence isolation, and escalate uncertainty?
02 · Evaluation benchmark
Built a disease-focused BioASQ-derived benchmark linking biomedical questions to reference PMIDs, answers, and evidence snippets, enabling consistent evaluation of retrieval quality and grounded synthesis across fixed test cases.
The versioned set contains 18 questions: 11 yes/no, 5 summary, 1 list, and 1 factoid. Twelve cases form the development split and six form the validation split.
Reference PMIDs, answers, documents, and snippets are used only during evaluation. They are not supplied to candidate retrieval or ranking.
| Split | Cases | Candidate retrieval | Ranking quality | ||
|---|---|---|---|---|---|
| Candidate recall | Recall@10 | nDCG@10 | MRR@10 | ||
| Development | 12 | 1.000 | 0.339 | 0.489 | 0.589 |
| Validation | 6 | 0.985 | 0.476 | 0.428 | 0.524 |
Measured limitation: nearly all reference publications entered the candidate pool, but fewer than half appeared in the top ten. Ranking—not candidate retrieval—is the primary measured limitation.
Candidate recall asks whether reference PMIDs entered the graph-derived candidate pool.
Recall@10 asks how much reference evidence appeared in the top ten.
nDCG@10 rewards reference publications placed nearer the top.
MRR@10 measures how early the first reference publication appeared.
18 schema-valid and provenance-valid records. Yes/no accuracy: 7 of 11 (0.636).
18 schema-valid and provenance-valid records. Yes/no accuracy: 8 of 11 (0.727).
The difference is descriptive and does not establish that retrieval caused an answer difference. Semantic and ROUGE similarity are not treated as factuality measures.
Retrieval metrics measure agreement with BioASQ references, not experimental directness, evidence strength, or scientific quality.
03 · Failure diagnosis
Localized failures to biological input validation, candidate retrieval, ranking, answer generation, citation checking, or automated judging so corrective actions could address the cause instead of relying on a single aggregate score.
Concrete finding: high candidate recall alongside weaker top-ten ranking isolates ranking as the measured bottleneck rather than graph candidate retrieval.
04 · Traceability
Linked retrieved publications and generated claims to their PMIDs, data and model versions, and processing steps. Added automated citation checks and Phoenix/OpenTelemetry traces across retrieval and answer generation.

05 · LLM-as-judge reliability
Tested an LLM-as-judge workflow and identified citation-confusion and low-label-diversity failures. Withheld unsupported quality claims pending expert review and judge calibration.
The judge found support elsewhere in the selected context rather than auditing the citation attached to the claim. This was recorded as an evaluator failure, not accepted as evidence of answer quality.