AI, Human Error and Consistency in Leukemia Research
Consistency is valuable, but repeating the same answer does not make that answer correct. AI research should examine both the stability of an output and the evidence that it is useful for its intended task.

Compare like with like
The AI-PAL study by Alcazer and colleagues developed and validated a model using routine laboratory parameters to predict acute leukemia subtypes. It is not a blood-smear image classifier. Its findings should therefore not be placed beside microscopy results as if the models performed an identical task. [1]
For a useful comparison, first identify the input, target population, reference standard and unit of analysis. Then examine whether the reported evaluation addresses the question being asked.
Reader agreement deserves its own study
Our recommendation is to measure disagreement directly. A study of inter-observer variability should describe who reviewed the cases, what instructions they received and how differences were resolved. A high model accuracy figure alone cannot show that AI reduced disagreement between people.
Similarly, a before-and-after claim should report the actual human workflow. Did reviewers see the original image? Could they reject the suggestion? Were difficult cases included? These details determine what the result means.
Keep the review record visible
A traceable research workflow should preserve the image, processing version, output, reviewer action and reason for any correction. This allows the team to distinguish a data issue from a model issue or an interpretation difference.
It also supports meaningful error review. A list of corrected cases is more useful when the team can reconstruct what happened and identify a recurring pattern.
Standardization is broader than automation
European LeukemiaNet's MRD consensus illustrates the importance of consistent technical methods and reporting in AML assessment. Automation is only one possible part of a workflow that still depends on defined procedures and interpretation. [2]
CellSight's perspective is to make evidence and review context accessible within the research record. This is a practical design priority, not a claim that the platform eliminates human error or has demonstrated a reduction in clinical disagreement.
References
- 1. Evaluation of a machine-learning model based on laboratory parameters for the prediction of acute leukaemia subtypes: a multicentre model development and validation study in France. Alcazer et al., The Lancet Digital Health. 2024. Accessed 2026-09-14.
- 2. 2021 Update on MRD in acute myeloid leukemia: a consensus document from the European LeukemiaNet MRD Working Party. Heuser et al., Blood. 2021-12-30. Accessed 2026-09-14.
Related Insights
Can AI Replace a Physician?
Explore the difference between automating a healthcare task and taking responsibility for care, with a practical perspective on human oversight of AI.
By CellSight Editorial Team · September 13, 2026 · 2 min read
Read articleHow Is AI Used in Healthcare?
Explore healthcare AI through clearly defined tasks, from image analysis to information retrieval, and learn why useful output still needs evidence and review.
By CellSight Editorial Team · September 13, 2026 · 2 min read
Read articleDesigning traceable AI image workflows
How identifiers, metadata and timestamps turn a microscopic image workflow into something that can be reviewed later.
Read article


