NoSQL databases, Hadoop, Big Data: Pinned tabs Nov.3rd
01: The real power of the Cassandra asynchronous database driver. The database is the heart, but its drivers are the veins. Just imagine what a clot would mean. â
(nb: if you decide to call me biased on this one, I challenge you to check the driverâs features and API and let me know what youâd like to see being done differently).
02: A (big) surprise for me to learn that Hortonworks and Databricks (Apache Spark) were not in a partnership. At least in the light of the work Hortonworks has already put into Spark. Now things are good again. â
03: Big name added to Aerospikeâs customer list: KAYAK. Deployed on 2 datacenters on real servers using SSDs, Aerospike replaced ehcache and memcache. â
04: After reading about Kayak and Aerospike, Iâve searched for similar case studies and found this article from Nov. 2013, about Amadeus and Couchbase. In that case the replaced technologies were Memcached and MySQL. â
05: Astute description of the current world of data: âIn a world where IT budgets grow at 2% per yearâat bestâand data is growing at 60% per year, something has to giveâ. â
06: Big data is (presented as) the Holy Grail for the sellers: Big Data analytics has the potential to allow sellers to directly influence buyers behavior based on their past behavior and habits. â
Iâd like to know whatâs the best a company can do for me (the customer). Unfortunately, I donât think there are that many companies where their customersâ and the company goals are aligned. For now, Iâll stick with my initial reaction: âman, this sort of future will really suckâ.
07: Hadoop would be perfect if only 12 things would suck about it. Not to mention that most of these are actually related to Oozie, Amabari, Knox, etc. I wholeheartedly agree though that a project with poor documentation and with strange and unexplainable exceptions and error messages will leave a very sour taste; so sour that many users will start looking for alternatives as soon as possible. â
08: Two numbers from Wikibonâs survey of the Hadoop market: 1) 25% of Hadoop users are paying for subscriptions; 2) 65% of Big Data practitioners1 have started migrating from Enterprise Datawarehouses to Hadoop and BigData. Very encouraging numbers. â
Other analysts have different data and thoughts though.
09: This article about Cassandra replacing Oracle at iovation is the perfect exemplification of the benefits of Cassandra: an always-on database, with better scalability and performance, development agility, all at a price that makes sense and scales with your business. Itâs also useful to see how the CTO of iovation has framed and prioritized the main issues he wanted addressed. â
10: A pretty good and well paced âlet me introduce you to HBaseâ including non-boring details about its roots, data model, installation and up to running some basic operations in the HBase shell. â
11: LinkedIn did some restructuring of their data science teams; a team which can be considered one of the first data science teams around (at least DJ Patil who lead this team before 2011 is said to have come up with the data scientist name). Many of the original members of the team have already left (and have now leading roles in their new places), but if your company needs an experienced data scientist, this might be a lucky time. â
12: If you like experimenting with data, Alexandre Passant has a blog post giving basic details about getting data from Deezer, uploading it to Google Cloud Storage and then using it from Googleâs BigQuery. â
Leaving aside the first step (getting the data), we have to agree that this process got a lot simpler over time and now itâs down to what ideas and algorithms you apply to the data.
Having data remains the big problem. That explains why companies fight and invest that much into different data platforms and try to get as many data points as possible. Get it now and youâll figure out how to use it later. The opposite wonât work.
13: Darshan of MongoDirector looks at the current trends related to databases, starting with DevOps and finishing with the numerous Database-as-a-Service products, to suggest that the future of DBAs is being part of the dev team as a developer specialized on databases. This is already quite true for small companies. Iâm not sure though about the enterprises that that revolve around governance and access rules and policies. â
14: A ton of papers about visualization categorized and presented in a very nice visualization. Even if you donât plan to read any of these papers, I strongly encourage you to take a look at this page to see how to build a helpful navigation and visualization of a swath of scientific papers. â
15: If you havenât read an article about the CAP theorem in a while, Robert Hodges wrote one in which he shares his thoughts about the possible differences between the theorem terms and real live, and then about its applicability in the real systems. Itâs not one of those posts proposing thereâs a way around the CAP theorem or the more radical the CAP theorem is wrong, buy our product. Itâs rather one in the vein of Daniel Abadiâs Replace CAP with PACELC: Problems with CAP, and Yahooâs little known NoSQL system. â
16: One of the major Hadoop events, Hadoop World, through the eyes of one of the top analysts, Merv Adrian from Gartner: last yearâs Hadoop 2.0 and YARN, SQL-on-HDFS noise, was replaced by Spark. And a lot of companies are jumping onboard. â
17: Mark Needham advises to use specific relationships (as opposed to generic ones) in Neo4j. And he presents the data behind his advice. â
18: A free mini-book, âThe DBAâs Guide to NoSQL and Apache Cassandraâ authored by a RDBMS DBA turned VP of Product2 at one of the leading NoSQL database providers, DataStax. â
Iâm not very sure how this is defined. â©
Heâs also my boss. And someone that I had and still have a lot to learn from. â©
Original title and link: NoSQL databases, Hadoop, Big Data: Pinned tabs Nov.3rd (NoSQL database©myNoSQL)