Welcome!

Recurring Revenue Authors: Yeshim Deniz, Liz McMillan, Carmen Gonzalez, Elizabeth White, Pat Romanski

Related Topics: @BigDataExpo, Java IoT, @CloudExpo

@BigDataExpo: Article

Data Lake Plumbers | @BigDataExpo @Schmarzo #BigData #IoT #AI #ML #DL

The data lake is ideal for your data science team as it liberates them from the constraints & limitations of the data warehouse

Many of my blogs promote the business benefits of the data lake, both from a “save me more money” as well as the “make me more money” perspectives. But I fear that I’m making this thing called the data lake sound like a “silver bullet" [1] – just drop the data into the data lake and everything magically works. But much like in the world of data warehousing, there is significant work that needs to be done under the covers – in areas such as metadata management, data governance and security – to ensure that the data lake will perform for a business in a production environment. Many of the processes and techniques we learned in the data warehouse will benefit us here, though there are many new tools to be aware of that can help us in the operationalization task.

I’ve asked an industry expert in metadata management and data governance, Joe DosSantos (follow Joe on twitter: @JoeDosSantos) to co-author this blog with me. Well, to be honest, this mostly reflects Joe’s experience and thinking; I just wanted to get credit for being smart enough to know when to bring someone smarter than me into the conversation!

Data Lake Benefits
You know from previous blogs that there are many benefits to the data lake including:

  • Capture data from wide range of traditional (operational, transactional) and new sources (structured and unstructured) as-is
  • Store all your data in one environment for cross-functional business analysis
  • Support the analytics and data science to uncover new customer, product, and operational insights
  • Empower front-line employees and managers, and drive a more profitable customer engagement leveraging customer, product and operational insights
  • Integrate analytic insights into operational (Finance, Manufacturing, Marketing, Sales Force, Procurement, Logistics) and management systems (Business Intelligence reports and dashboards)

The data lake is ideal for your data science team in that it liberates them from the constraints and limitations of the data warehouse, enabling the data science team to quickly ingest, test and determine if there is any value to different data sets and analytic techniques without having to go through the rigorous operational procedures that govern the data warehouse.

However, this liberty can be quite terrifying in highly regulated environments. Companies have spent years developing governance and stewardship organizations specifically to control patient information, personal contact information, account balances, and other sensitive information. The description above seems to undo all of this work by creating free and easy access to data that should be locked down.

This is why the controls of a data lake need to be very clear. Data that is onboarded into a lake must go through a rigorous set of operational procedures to manage and govern that data set to make sure that it is appropriately tagged and protected, and then provisioned only to people who have the proper authorization. Modern data tools allow for this kind of governance capability to balance the quick and easy access to data that a data scientist needs with the security that good practices (and often the government) demand.

Operationalizing the Data Lake
Operationalizing the data lake requires several non-obvious disciplines, many of which we learned from our data warehouse experiences. These disciplines include data ingestion, indexing, cataloging, metadata management, data governance and security [2].

  1. As with a data warehouse, you will need a method to bring data into your environment. As batch windows became longer and longer in the data warehouse world and business users clamored for increasingly up-to-date information, practitioners began shifting from conventional data ETL (Extract, Transform, Load) to lower latency streaming and micro-batch. This trend was extended in the big data universe with tools like Kafka, a streaming message bus, and with Spring and Sqoop to accelerate data ingest. In the big data world, you might also want to ingest unstructured data sets as well, introducing new tools like Flume. Finally, you may want to trigger complex events based on this data stream and you might do so via Spark, Gemfire, or other in-memory grids. And just to make it more complex, you will likely use several of these tools in combination depending on your data feed needs. Keep in mind that in the world of ELT (Extract, Load, Transform) (note that the order differs from E-T-L), most of these data movements are data dumps. At this point, you have simply collected lots of raw data. It’s now time to make sense of it.
  2. Next, it is useful to tag files that you have ingested. What kind of file is this? What would be useful to know about it so that I could search for it later? Zaloni Bedrock is an example of a tool to apply metadata tags to the files, which is useful for both structured and unstructured data sets.
  3. We mentioned above that one of the key requirements of our data lake is having control over who can have access to specific data sources. Generally speaking, the data loaded in steps 1 and 2 is what we call “Bronze” data, meaning that it is good enough for the business process that created it. Data in these sets will likely be sensitive and your security should reflect it. However, we need to determine rules for how the data should be modified, obfuscated, or deleted in order to make it consumable for broader audience, or what we might call “Silver” status. You need to create business rules to manage data (e.g. birthdays should be masked and social security numbers should be stored as only the last 4). Collibra is an example of a tool for this rules definition and management. It allows data rules to be set up based on logical business entities by business people rather than technologists.
  4. For those people who are familiar with governance concepts, you will recognize the difference between a policy and a control. A policy is like a speed limit sign along the highway. The control is the police officer that pulls you over if you are driving over that speed limit. Data works the same way. While Collibra establishes the policy, you need to create a method for enforcing that policy. To do this, you need to find the logical entities buried in the data (i.e. “oh look, I found a social security number!”). Examples of such products include Global IDs for scanning structured data sets with the operational systems and Waterline for scanning data inside of Hadoop.
  5. Once you have found the data that you want, you want to implement the rules. For this, there is an open source tool called Atlas that contains an orchestration capability called Falcon that helps implement the rules.
    1. Apache Atlas is a scalable and extensible set of core foundational governance services that enables enterprises to effectively and efficiently meet their compliance requirements within Hadoop and allows integration with the complete enterprise data ecosystem.
    2. Apache Falcon is a data governance engine that defines, schedules, and monitors data management policies. Falcon allows Hadoop administrators to centrally define their data pipelines, and then Falcon uses those definitions to auto-generate workflows in Apache Oozie
  6. Now that the data is loaded, you will want to enforce security through your LDAP capability or possibly through Kerberos. There are also tools like Blue Talon that simplify the ability to authorize, provision, protect, enforce and audit data security policies across your data lake.
  7. Finally, audit controls are critical. Cloudera introduced Navigator specifically to allow simple transparency to data history and lineage. Hortonworks will rely on Atlas to provide this capability.

Data that has gone through the above processes creates a view and accessibility of the data that can be made available to a wide set of users – both business analysts and data science teams.

Summary
When you build a house, the vast majority of the creative work is in the features and curbside appeal. That’s the fun part. But without the underlying plumbing, the house would quickly degrade into a money pit.

Consider the metaphor of a retail store: stocking the shelves vs. purchasing goods. When you go to the store, you don’t care about how the goods got there, but the rules for accessing the goods are everywhere. Cigarettes are behind the front desk. Pharmaceuticals must be dispensed with a prescription. Razor blades are under lock and key (for some strange reason). There are policies and enforcements on stocking the shelves so that the shopping experience is clear and easy.

To successfully operationalize the data lake, organizations need to address all of the plumbing requirements outlined in this blog that enable the business users and data science teams to have confidence in the wealth of data that the organization is amassing. The data lake plumbing processes may not be very glamorous, but without them, you’ll end up with a stinky data dump instead of a glorious data lake.

References:

  1. A “silver bullet” is a simple and seemingly magical solution to a complicated problem.
  2. While I mention several tools, this blog is not meant to be an endorsement of these tools nor is this intended to be a comprehensive list of such tools. However, many of these tools are the same tools that we use in our data lake implementations at EMC.

Data Lake Plumbers: Operationalizing the Data Lake
Bill Schmarzo

More Stories By William Schmarzo

Bill Schmarzo, author of “Big Data: Understanding How Data Powers Big Business”, is responsible for setting the strategy and defining the Big Data service line offerings and capabilities for the EMC Global Services organization. As part of Bill’s CTO charter, he is responsible for working with organizations to help them identify where and how to start their big data journeys. He’s written several white papers, avid blogger and is a frequent speaker on the use of Big Data and advanced analytics to power organization’s key business initiatives. He also teaches the “Big Data MBA” at the University of San Francisco School of Management.

Bill has nearly three decades of experience in data warehousing, BI and analytics. Bill authored EMC’s Vision Workshop methodology that links an organization’s strategic business initiatives with their supporting data and analytic requirements, and co-authored with Ralph Kimball a series of articles on analytic applications. Bill has served on The Data Warehouse Institute’s faculty as the head of the analytic applications curriculum.

Previously, Bill was the Vice President of Advertiser Analytics at Yahoo and the Vice President of Analytic Applications at Business Objects.

@ThingsExpo Stories
While the focus and objectives of IoT initiatives are many and diverse, they all share a few common attributes, and one of those is the network. Commonly, that network includes the Internet, over which there isn't any real control for performance and availability. Or is there? The current state of the art for Big Data analytics, as applied to network telemetry, offers new opportunities for improving and assuring operational integrity. In his session at @ThingsExpo, Jim Frey, Vice President of S...
DX World EXPO, LLC., a Lighthouse Point, Florida-based startup trade show producer and the creator of "DXWorldEXPO® - Digital Transformation Conference & Expo" has announced its executive management team. The team is headed by Levent Selamoglu, who has been named CEO. "Now is the time for a truly global DX event, to bring together the leading minds from the technology world in a conversation about Digital Transformation," he said in making the announcement.
"We provide IoT solutions. We provide the most compatible solutions for many applications. Our solutions are industry agnostic and also protocol agnostic," explained Richard Han, Head of Sales and Marketing and Engineering at Systena America, in this SYS-CON.tv interview at @ThingsExpo, held June 6-8, 2017, at the Javits Center in New York City, NY.
"We are focused on SAP running in the clouds, to make this super easy because we believe in the tremendous value of those powerful worlds - SAP and the cloud," explained Frank Stienhans, CTO of Ocean9, Inc., in this SYS-CON.tv interview at 20th Cloud Expo, held June 6-8, 2017, at the Javits Center in New York City, NY.
"DX encompasses the continuing technology revolution, and is addressing society's most important issues throughout the entire $78 trillion 21st-century global economy," said Roger Strukhoff, Conference Chair. "DX World Expo has organized these issues along 10 tracks with more than 150 of the world's top speakers coming to Istanbul to help change the world."
"We've been engaging with a lot of customers including Panasonic, we've been involved with Cisco and now we're working with the U.S. government - the Department of Homeland Security," explained Peter Jung, Chief Product Officer at Pulzze Systems, in this SYS-CON.tv interview at @ThingsExpo, held June 6-8, 2017, at the Javits Center in New York City, NY.
With tough new regulations coming to Europe on data privacy in May 2018, Calligo will explain why in reality the effect is global and transforms how you consider critical data. EU GDPR fundamentally rewrites the rules for cloud, Big Data and IoT. In his session at 21st Cloud Expo, Adam Ryan, Vice President and General Manager EMEA at Calligo, will examine the regulations and provide insight on how it affects technology, challenges the established rules and will usher in new levels of diligence...
The financial services market is one of the most data-driven industries in the world, yet it’s bogged down by legacy CPU technologies that simply can’t keep up with the task of querying and visualizing billions of records. In his session at 20th Cloud Expo, Karthik Lalithraj, a Principal Solutions Architect at Kinetica, discussed how the advent of advanced in-database analytics on the GPU makes it possible to run sophisticated data science workloads on the same database that is housing the rich...
SYS-CON Events announced today that Massive Networks will exhibit at SYS-CON's 21st International Cloud Expo®, which will take place on Oct 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. Massive Networks mission is simple. To help your business operate seamlessly with fast, reliable, and secure internet and network solutions. Improve your customer's experience with outstanding connections to your cloud.
Internet of @ThingsExpo, taking place October 31 - November 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA, is co-located with 21st Cloud Expo and will feature technical sessions from a rock star conference faculty and the leading industry players in the world. The Internet of Things (IoT) is the most profound change in personal and enterprise IT since the creation of the Worldwide Web more than 20 years ago. All major researchers estimate there will be tens of billions devic...
"The Striim platform is a full end-to-end streaming integration and analytics platform that is middleware that covers a lot of different use cases," explained Steve Wilkes, Founder and CTO at Striim, in this SYS-CON.tv interview at 20th Cloud Expo, held June 6-8, 2017, at the Javits Center in New York City, NY.
SYS-CON Events announced today that Calligo, an innovative cloud service provider offering mid-sized companies the highest levels of data privacy and security, has been named "Bronze Sponsor" of SYS-CON's 21st International Cloud Expo ®, which will take place on Oct 31 - Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. Calligo offers unparalleled application performance guarantees, commercial flexibility and a personalised support service from its globally located cloud plat...
Everything run by electricity will eventually be connected to the Internet. Get ahead of the Internet of Things revolution and join Akvelon expert and IoT industry leader, Sergey Grebnov, in his session at @ThingsExpo, for an educational dive into the world of managing your home, workplace and all the devices they contain with the power of machine-based AI and intelligent Bot services for a completely streamlined experience.
SYS-CON Events announced today that DXWorldExpo has been named “Global Sponsor” of SYS-CON's 21st International Cloud Expo, which will take place on Oct 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. Digital Transformation is the key issue driving the global enterprise IT business. Digital Transformation is most prominent among Global 2000 enterprises and government institutions.
SYS-CON Events announced today that Datera, that offers a radically new data management architecture, has been named "Exhibitor" of SYS-CON's 21st International Cloud Expo ®, which will take place on Oct 31 - Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. Datera is transforming the traditional datacenter model through modern cloud simplicity. The technology industry is at another major inflection point. The rise of mobile, the Internet of Things, data storage and Big...
"MobiDev is a Ukraine-based software development company. We do mobile development, and we're specialists in that. But we do full stack software development for entrepreneurs, for emerging companies, and for enterprise ventures," explained Alan Winters, U.S. Head of Business Development at MobiDev, in this SYS-CON.tv interview at 20th Cloud Expo, held June 6-8, 2017, at the Javits Center in New York City, NY.
SYS-CON Events announced today that DXWorldExpo has been named “Global Sponsor” of SYS-CON's 21st International Cloud Expo, which will take place on Oct 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. Digital Transformation is the key issue driving the global enterprise IT business. Digital Transformation is most prominent among Global 2000 enterprises and government institutions.
In his opening keynote at 20th Cloud Expo, Michael Maximilien, Research Scientist, Architect, and Engineer at IBM, discussed the full potential of the cloud and social data requires artificial intelligence. By mixing Cloud Foundry and the rich set of Watson services, IBM's Bluemix is the best cloud operating system for enterprises today, providing rapid development and deployment of applications that can take advantage of the rich catalog of Watson services to help drive insights from the vast t...
SYS-CON Events announced today that EnterpriseTech has been named “Media Sponsor” of SYS-CON's 21st International Cloud Expo, which will take place on Oct 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. EnterpriseTech is a professional resource for news and intelligence covering the migration of high-end technologies into the enterprise and business-IT industry, with a special focus on high-tech solutions in new product development, workload management, increased effic...
SYS-CON Events announced today that Massive Networks, that helps your business operate seamlessly with fast, reliable, and secure internet and network solutions, has been named "Exhibitor" of SYS-CON's 21st International Cloud Expo ®, which will take place on Oct 31 - Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. As a premier telecommunications provider, Massive Networks is headquartered out of Louisville, Colorado. With years of experience under their belt, their team of...