NHS SynAE liveblog
This is a static transcript of a 24liveblog
Last Updated: 05/03/2019 15:40
05/03/2019 15:40 Wrap up and next steps
Wrapping up the day, the general consensus in the room was that synthetic data was exciting but was very much unknown territory. There is the potential to be pioneers for this type of data and to set the standards and definition that others could follow, but there's also a huge amount of risk attached to being adventurous.
There was a collective agreement that there was more risk of utility than privacy, meaning there was probably more risk of the data not being used or being unfit for purpose than there being personal data problems.
The SynAE project continues, with another potential data sample release later in the year. ODI Leeds will continue to work with ODI HQ and NHS England as we go forward, so stay tuned for project updates and more opportunities to get involved :)
Thank you for joining us today! For all of the resources relating to today, head over to the SynAE website.
05/03/2019 15:27 Use cases and data format
Just after lunch, a sample dataset was shared with the attendees (not available outside of the workshop). This sample dataset was 10,000 lines from a random point in the source synthetic dataset that Jonny has been working on, so records have been anonymised and swapped.
Following from Tom's spectrum, Paul Connell then lead a discussion about what synthetic data will be used for. If there is so much uncertainty around synthetic data, from the privacy concerns to the utility of it, then what's the point?
- To better design public services
- Procurement - accelerate the service design and build process. Less personas more data - but not real fewer models more measured outcomes.
- Open development approach.
- To prove a method
- To test a model
- To create an eco-system
- Governance
- To find new and innovative solutions.
- Accelerate the publishing of data and Streamlining the process to get to real data.
- Imagine new questions to ask of the data.
- Combine data sources to create
- Test transparency options.
- Discoverability within The NHS
05/03/2019 15:02 Tom's synthetic spectrum
In an attempt to answer several of the questions posed before lunch, Tom Forth (Head of Data for ODI Leeds) embarked on a mapping mission. Taking over the huge whiteboard wall in the main space at ODI Leeds, he began thinking out loud with the different types of synthetic data release and the associated political/safety concerns.
For example, at one end of the scale we have 'Say there is no data at all', followed by 'Admit there is data but don't release anything.' At the opposite end of the scale, we have 'Release everything' or 'Release the de-identified raw data' as the lesser extreme. Both have very defined reactions, both politically and in terms of personal safety.
Through this exploration, Tom and the group worked towards a consensus that synthetic data is better than just a schema, and a bit more detailed synthetic data is better than simplified synthetic data. But whatever the type of synthetic data, there will always be some kind of trade-off between privacy and utility.
05/03/2019 12:31 Just before lunch...
Tantalisingly close to lunch time, Paul Connell then asked the room to work up some potential challenges and themes to investigate in the afternoon.
- testing
- technical
- SynAE vs schema
- procurement
- what is the point?
- when to use synthetic data vs other org techniques?
- data is about relationships - can this be replicated?
- at what point is data personal?
- is data personally identifiable because we are so free with our benign data?
- why is SynAE better than a schema?
05/03/2019 11:58 Questions from the room on risk
Good comment from the audience - given the pace of technological advances, modelling for threats and risk can be really hard. A few years ago, using Google to fact-check or identify people from news stories was not possible.
Question about risks - if the fear of releasing data means that it becomes so diluted as to be useless, surely that is also a threat and risk too? What good is a data release if no one can use it?
Did you get into the opportunity mapping as well? Well, no. It was partly saved for the SynAE today.
05/03/2019 11:44 Risky questions
The main questions that Fiontánn and Olivier used in their threat modelling approach:
- if I already know someone is in this dataset, can I find them?
- if I already know someone is in this dataset, what more can I learn about them?
There was also consideration given to the victim - what could they do that increases the risk of re-identification or what could they do to decrease the risk?
Aside from the threats to people, they identified other areas that could create different kinds of threats and risks:
- false insight: poor decisions made from the data
- anonymisation process: people treating the data as real, even if it isn't
- fear: too risky, too much complication?
05/03/2019 11:34 ODI HQ - using tools for anonymisation threat modelling
Up next is Oliver and Fiontánn from the Open Data Institute, who had embarked on a project about anonymisation and tools to help just as the SynAE project was in development. There was a lot of crossover between what Olivier and Fiontánn were exploring and what the SynAE project would encounter.
In relation to SynAE, Olivier and Fiontánn focused especially on mapping the data ecosytem and threat modelling, which would complement the discussions today looking at the challenges to the methodology.
The ecosystem mapping was all about identifying sources of data, beneficiaries, custodians, stewards, gateways, etc. All of the players and paths for the data in the SynAE project.
05/03/2019 11:18 Questions from the audience
By grouping before variable swapping, are we building bias into the dataset? For example, using IMD (indices of multiple deprivation) can reveal patients that all come from similar socio-economic backgrounds. How could this have an impact on healthcare decisions? For good, or for bad?
Perhaps using precedents would help choose the most appropriate variable swaps. For instance, for location, using the LSOA rather than the IMD.
05/03/2019 11:09 This definition of synthetic data
Once the anonymisation process is complete, that provides the format for the data and provides confidence that a single record could not be re-identified. At this point, Jonny explored the need to remove original data from the variables that remain. He applied techniques like a 'random forest' or Pearson's correlation to find the variables that could be swapped and retain the usefulness. These are then swapped.
This is the chosen definition of synthetic data for SynAE. The data is tested for duplication, errors, retainment, etc.
05/03/2019 10:54 Granularity versus useful insight
Jonny takes us through the process that he has been working to, which uses cumulative cycles of anonymisation to strip away identifying data and then alter what remains so that an individual cannot be identified.
At the point of talking about the grouping of age, a member of the audience had a question. He is currently working with data about DNAs (did not attend) within the NHS. Some of their analysis has been looking at specific trends within age, so would the synthetic data be granular enough?
05/03/2019 10:39 The process of synthetic data
Jonny Pearson, Senior Analytical Manager for NHS England, takes us through his approach to creating synthetic data. He begins with a similar attitude to Tom - he wants to know what people want to use the data for, what's missing, what's wrong/right, etc. For him, there are missed opportunities to find patterns that could lead to wider healthcare strategy and prevention.
Use cases that they've identified so far are:
- NHS analysts
- commercial entities
- academic
- innovation
- other synthetic data developers
The method that Jonny has used begins with anonymisation. Identifying data is stripped away or replaced with a different value, for example a name replaced with a random patient ID number. He continued to move through each variable and remove granularity but still retain usefulness.
Jonny wrote a blog post ahead of today's event that contains more detail.
05/03/2019 10:25 Millions attend A&E every year
Tom Forth quickly summarises the how and why of focusing on A&E data specifically. With millions of visits to A&E every year, the only data that the public - and many internal analysts - get to see is an aggregated and summary version. Which means we are losing a vast amount of vital data that could be used to bring benefits to the NHS and to patients, even to wider healthcare strategy.
For Tom, he wants the following to happen today:
- share ideas of what could be made with this data
- tell us what is good or wrong about the sample dataset
- suggestions for the dataset going forward
05/03/2019 10:17 How we do things...
Our method for helping people is to make things open. We work with them to define what they want to do or the questions they want to ask, develop 'the why', and then get people involved. We encourage healthy challenge and debate around a project as this can help shape opportunities.
05/03/2019 10:14 Why are we here?
Not an existential question, we promise. Today's event ultimately came out of the first discussions between Paul Connell and Forrest at NHS England, who
Ming from NHS England DAIS is excited about today. A big challenge for the NHS has been safely publishing data. They face barriers around information governance and regularly get asked about 'use cases' to justify the release of data. So the SynAE event is an opportunity to explore methods that will allow for datasets to be published where sensitive information has been stripped away or replaced. All of this also relates to developing the transparency of the NHS.
05/03/2019 10:07 Kicking off
We're just allowing a bit of time for folk to arrive but in the mean time, Paul Connell and Tom Forth of ODI Leeds are just making gentle introductions :) ODI Leeds was founded with the intention of exploring and delivering value from open innovation and open data that benefits everyone.


