Is Big Data Safe?
People keep talking about big data and how important it is. They talk about how useful it is in making predictions and infographics, or how much it can be worth. But how does big data affect individuals like you and me?
Big Data: Data sets containing much larger amounts of data than what would normally be processed. It can be in the terabytes or up to zettabytes of information gathered from many sources and about many things. It is usually defined by the volume, velocity and variety of the data.
Big data has many uses. It can be used by companies like Netflix to anticipate trends based on user activity. They look at things like how people have acted in the past, what affected how they acted and what is happening now. It's also used to find flaws in things by seeing where people struggle with them and where they tend to fail. Big data is used to train Artificial Intelligence (AI), and to maximise internal operational efficiency. According to the UN Environment Programme, big data can help assess environmental risks and minimise the harm done to it.
According to IBM, Artificial intelligence leverages computers and machines to mimic the problem-solving and decision-making capabilities of the human mind. That means that essentially all it has to be is an algorithm that makes decisions to qualify as AI. AI can then be "weak" or "strong", where weak ones have a single or narrow purpose. For an example, consider Deep Blue, the chess computer.
Like every other seemingly perfect solution, big data isn't without its risks. First we have to as ourselves how exactly they are getting this big data. Such huge amounts of information have to come from somewhere, and that place is usually from us. Most things you do can be and often are tracked. If you go to websites, not only will they, Google and often Facebook see that you are there and everything you do there, but often they will have additional cookies track you across the internet afterwards. If you use the Wi-Fi at a café you visit, or ask Alexa about the weather in Paris, that may be saved. These reams of data can then be used to try to answer any questions companies may have about you, like how you would respond to specific adverts or if you are likely to buy something.
The General Data Protection Regulation (GDPR), a fairly recent EU regulation that controls data use, has as some of its core principles Data Minimisation and Purpose Limitation. The former is the idea that companies should only have the information they absolutely need to do their job. It clashes quite directly with big data, which necessitates having as much data as possible. The latter also clashes because it is to only use data for a set, closed purpose whereas big data gathers the data first and uses it to answer questions later.
Things get even more complicated when you start thinking about the kind of AI that learns on its own. They use big data to change themselves, and then make more decisions based on big data. This makes them minimally transparent, and increases the risk of bias. Especially if the data it is learning from is skewed, its predictions may quickly become discriminatory. For example, an algorithm trained on the massively Black and Latin NYPD gang ties database will almost certainly flag up far more minorities than white people in the future. Also, an algorithm being used in healthcare to sort people by how much help they needed gave equally sick black people lower risk scores because of the specific terms it used.
Can the AI making decisions based on big data ever really be accountable either? Accountability is another GDPR principle, but it becomes complicated when the processing is automatic and at such a massive scale. Who exactly is responsible for the results of the processing? Who is responsible for how the data was acquired? How do people know what data is being processed and what it is being processed for? The ICO also says that for an AI's processing to be accountable, bias has to be avoided before it affects people rather than spotted after. How do we do that?
The ICO is the Information Commissioner's Office, the entity whose role is to ensure data protection and that the GDPR is followed. They are the ones you can complain to if you feel that your data rights are being infringed.
All in all, while big data has a lot of potential, it also has many risks. If we can ensure that the data does not relate to humans when it does not have to (such as when it's about the environment) and that when it does the specifics are clear to every subject and they have given their consent then big data has the potential to make the world a better, more efficient place. However, we need to be very careful that the price of efficiency is not our data safety.
('AI and data protection: balancing tensions' by Rob Sumroy and Natalie Donovan on Thomson Reuters Practical Law)















