Resource Origin: DataFloq
“It is generally accepted that big data can be explained according to three V’s: Velocity, Variety and Volume. In a 2001 research report, META Group (now Gartner) analyst Doug Laney defined big data as being three-dimensional, i.e. increasing volume (amount of data), velocity (speed of data in and out), and variety (range of data types and sources). Later in 2012 Gartner updated the definition of big data as high volume, high velocity, high variety.
Although I do not want to diminish the importance of the definition by Gartner, I do think that big data can be better explained by adding a few more V’s. These V’s explain important aspects of big data and a big data strategy that organisation cannot ignore. Let’s look at these V’s and for completeness, let’s also once more mention the common known V’s:
Velocity
The Velocity is the speed at which the data is created, stored, analysed and visualised. In the past, when batch processing was common practice, it was normal to receive an update from the database every night or even every week. Computers and servers required substantial time to process the data and update the databases. In the big data era, data is created in real-time or near real-time. With the availability of Internet-connected devices, wireless or wired, machines and devices can pass-on their data the moment it is created.
The speed at which data is created currently is almost unimaginable: Every minute we upload 100 hours of video on Youtube. In addition, every minute over 200 million emails are sent, around 20 million photos are viewed and 30.000 uploaded on Flickr, almost 300.000 tweets are sent, and almost 2,5 million queries on Google are performed. The challenge organisations have to cope with the enormous speed the data is created and used in real-time.
Volume
90% of all data ever created, was created in the past two years. From now on, the amount of data in the world will double every two years. By 2020, we will have 50 times the amount of data as that we had in 2011. The sheer volume of the data is enormous, and a very large contributor to the ever-expanding digital universe is the Internet of Things with sensors all over the world in all devices creating data every second. The era of a trillion sensors is upon us. If we look at aeroplanes they generate approximately 2,5 billion Terabyte of data each year from the sensors installed in the engines. Self-driving cars will generate 2 Petabyte of data every year.
Also, the agricultural industry generates massive amounts of data with sensors installed in tractors. Shell uses super-sensitive sensors to find additional oil in wells, and if they install these sensors at all 10.000 wells, they will collect approximately 10 Exabyte of data annually. That again is absolutely nothing if we compare it to the Square Kilometer Array Telescope that will generate 1 Exabyte of data per day. In the past, the creation of so much data would have caused serious problems. Nowadays, with decreasing storage costs, better storage solutions like Hadoop and the algorithms to create meaning from all that data this is not a problem at all.
Variety
In the past, all data that was created was structured data, it neatly fitted in columns and rows but those days are over. Nowadays, 90% of the data that is generated by the organisation is unstructured data. Data today comes in many different formats: structured data, semi-structured data, unstructured data and even complex structured data.
The wide variety of data requires a different approach as well as different techniques to store all raw data. There are many different types of data, and each of those types of data requires different types of analyses or different tools to use. Social media like Facebook posts or Tweets can give different insights, such as sentiment analysis on your brand, while sensory data will give you information about how a product is used and what the mistakes are…”
https://datafloq.com/read/3vs-sufficient-describe-big-data/166
Contributors: Open SourceCategories: SIG U Resource, Tool
SRC Type: Disruptive Digitization, Machine Learning