Skip to main content
All insights
Responsible AI

AI, Human Error and Consistency in Leukemia Research

Consistency is valuable, but repeating the same answer does not make that answer correct. AI research should examine both the stability of an output and the evidence that it is useful for its intended task.

By CellSight Editorial TeamOriginally published Updated 2 min read
Conceptual comparison workspace with two optical eyepieces, three glass slides and a violet notebook.
Original AI-generated editorial illustration; not a clinical image, result, or photograph of CellSight facilities.

Compare like with like

The AI-PAL study by Alcazer and colleagues developed and validated a model using routine laboratory parameters to predict acute leukemia subtypes. It is not a blood-smear image classifier. Its findings should therefore not be placed beside microscopy results as if the models performed an identical task. [1]

For a useful comparison, first identify the input, target population, reference standard and unit of analysis. Then examine whether the reported evaluation addresses the question being asked.

Reader agreement deserves its own study

Our recommendation is to measure disagreement directly. A study of inter-observer variability should describe who reviewed the cases, what instructions they received and how differences were resolved. A high model accuracy figure alone cannot show that AI reduced disagreement between people.

Similarly, a before-and-after claim should report the actual human workflow. Did reviewers see the original image? Could they reject the suggestion? Were difficult cases included? These details determine what the result means.

Keep the review record visible

A traceable research workflow should preserve the image, processing version, output, reviewer action and reason for any correction. This allows the team to distinguish a data issue from a model issue or an interpretation difference.

It also supports meaningful error review. A list of corrected cases is more useful when the team can reconstruct what happened and identify a recurring pattern.

Standardization is broader than automation

European LeukemiaNet's MRD consensus illustrates the importance of consistent technical methods and reporting in AML assessment. Automation is only one possible part of a workflow that still depends on defined procedures and interpretation. [2]

CellSight's perspective is to make evidence and review context accessible within the research record. This is a practical design priority, not a claim that the platform eliminates human error or has demonstrated a reduction in clinical disagreement.

References

  1. 1. Evaluation of a machine-learning model based on laboratory parameters for the prediction of acute leukaemia subtypes: a multicentre model development and validation study in France. Alcazer et al., The Lancet Digital Health. 2024. Accessed 2026-09-14.
  2. 2. 2021 Update on MRD in acute myeloid leukemia: a consensus document from the European LeukemiaNet MRD Working Party. Heuser et al., Blood. 2021-12-30. Accessed 2026-09-14.

Related Insights

Responsible AI

Can AI Replace a Physician?

Explore the difference between automating a healthcare task and taking responsibility for care, with a practical perspective on human oversight of AI.

By CellSight Editorial Team · September 13, 2026 · 2 min read

Read article
Responsible AI

How Is AI Used in Healthcare?

Explore healthcare AI through clearly defined tasks, from image analysis to information retrieval, and learn why useful output still needs evidence and review.

By CellSight Editorial Team · September 13, 2026 · 2 min read

Read article