Using routine healthcare data without its context may lead to misleading conclusions

July 27, 2026

A paper from researchers and clinicians from the Improvement Academy demonstrates that including routine data from electronic patient records in research without understanding the context in which it was collected may lead to inaccurate conclusions. Published in BMJ Health and Care Informatics in May, the authors suggest working closely with local data providers is crucial to get clarity on how the dataset was created before using it in research.

Patient data from routine appointments and treatments – including blood results, length of admission, resuscitation status, and diagnoses – are now widely available in anonymised form to researchers via regional Secure Data Environments (SDEs). However, aggregating data into ‘data lakes’ without an understanding of the local processes that were used to collect that data can lead to misleading results.

Lead author Dr Grace Duffy, Clinical Fellow at the Improvement Academy and Palliative Medicine Consultant, explained: “Routine data are collected by hospitals and other healthcare settings to help provide and manage patient care. Unlike research data, which are designed to answer specific research questions and record facts, routine data record processes. These processes may vary between locations, data inputters and timeframes, which is why it’s vital for researchers to understand the processes being measured in order to draw accurate conclusions from routine data.”

Co-author Dr Vishal Sharma, Connected West Yorkshire Co-Delivery Lead, Associate Director of the Improvement Academy, and Implementation Co-Lead for the Yorkshire and Humber NIHR Applied Research Collaboration, added: “Differences in the way data is recorded can have a significant impact on the story that data is telling us. For example, differences in how two hospitals record activity can make comparisons between those two hospitals difficult if the underlying processes are not understood.”

The paper includes a number of real-world scenarios where the context in which routine data were collected has caused problems for researchers, for example:

  • Different processes being used when a hospital changed from paper records to Electronic Health Records (EHR), leading to erroneous spikes in admission numbers
  • Differences in the SNOMED codes used by different NHS Trusts to indicate a patient has died in hospital, leading to one Trust showing lower than expected rates
  • Patients not routinely having their ethnicity data recorded, making it harder to show care was being provided equitably across the community.

The authors recommend that researchers should engage with the local teams who produce the routine data to create a clear understanding of which processes they measure, how they are collected and to consider together whether changes in the data reflect a change in process or a change in reality.

The paper is freely available to read in BMJ Health and Care Informatics.

Privacy

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.

For more information, please visit our Privacy Policy.