Welcome!

Recurring Revenue Authors: Yeshim Deniz, Pat Romanski, Stackify Blog, William Schmarzo, Carmen Gonzalez

Related Topics: @BigDataExpo, Java IoT, @CloudExpo

@BigDataExpo: Article

Data Lake Plumbers | @BigDataExpo @Schmarzo #BigData #IoT #AI #ML #DL

The data lake is ideal for your data science team as it liberates them from the constraints & limitations of the data warehouse

Many of my blogs promote the business benefits of the data lake, both from a “save me more money” as well as the “make me more money” perspectives. But I fear that I’m making this thing called the data lake sound like a “silver bullet" [1] – just drop the data into the data lake and everything magically works. But much like in the world of data warehousing, there is significant work that needs to be done under the covers – in areas such as metadata management, data governance and security – to ensure that the data lake will perform for a business in a production environment. Many of the processes and techniques we learned in the data warehouse will benefit us here, though there are many new tools to be aware of that can help us in the operationalization task.

I’ve asked an industry expert in metadata management and data governance, Joe DosSantos (follow Joe on twitter: @JoeDosSantos) to co-author this blog with me. Well, to be honest, this mostly reflects Joe’s experience and thinking; I just wanted to get credit for being smart enough to know when to bring someone smarter than me into the conversation!

Data Lake Benefits
You know from previous blogs that there are many benefits to the data lake including:

  • Capture data from wide range of traditional (operational, transactional) and new sources (structured and unstructured) as-is
  • Store all your data in one environment for cross-functional business analysis
  • Support the analytics and data science to uncover new customer, product, and operational insights
  • Empower front-line employees and managers, and drive a more profitable customer engagement leveraging customer, product and operational insights
  • Integrate analytic insights into operational (Finance, Manufacturing, Marketing, Sales Force, Procurement, Logistics) and management systems (Business Intelligence reports and dashboards)

The data lake is ideal for your data science team in that it liberates them from the constraints and limitations of the data warehouse, enabling the data science team to quickly ingest, test and determine if there is any value to different data sets and analytic techniques without having to go through the rigorous operational procedures that govern the data warehouse.

However, this liberty can be quite terrifying in highly regulated environments. Companies have spent years developing governance and stewardship organizations specifically to control patient information, personal contact information, account balances, and other sensitive information. The description above seems to undo all of this work by creating free and easy access to data that should be locked down.

This is why the controls of a data lake need to be very clear. Data that is onboarded into a lake must go through a rigorous set of operational procedures to manage and govern that data set to make sure that it is appropriately tagged and protected, and then provisioned only to people who have the proper authorization. Modern data tools allow for this kind of governance capability to balance the quick and easy access to data that a data scientist needs with the security that good practices (and often the government) demand.

Operationalizing the Data Lake
Operationalizing the data lake requires several non-obvious disciplines, many of which we learned from our data warehouse experiences. These disciplines include data ingestion, indexing, cataloging, metadata management, data governance and security [2].

  1. As with a data warehouse, you will need a method to bring data into your environment. As batch windows became longer and longer in the data warehouse world and business users clamored for increasingly up-to-date information, practitioners began shifting from conventional data ETL (Extract, Transform, Load) to lower latency streaming and micro-batch. This trend was extended in the big data universe with tools like Kafka, a streaming message bus, and with Spring and Sqoop to accelerate data ingest. In the big data world, you might also want to ingest unstructured data sets as well, introducing new tools like Flume. Finally, you may want to trigger complex events based on this data stream and you might do so via Spark, Gemfire, or other in-memory grids. And just to make it more complex, you will likely use several of these tools in combination depending on your data feed needs. Keep in mind that in the world of ELT (Extract, Load, Transform) (note that the order differs from E-T-L), most of these data movements are data dumps. At this point, you have simply collected lots of raw data. It’s now time to make sense of it.
  2. Next, it is useful to tag files that you have ingested. What kind of file is this? What would be useful to know about it so that I could search for it later? Zaloni Bedrock is an example of a tool to apply metadata tags to the files, which is useful for both structured and unstructured data sets.
  3. We mentioned above that one of the key requirements of our data lake is having control over who can have access to specific data sources. Generally speaking, the data loaded in steps 1 and 2 is what we call “Bronze” data, meaning that it is good enough for the business process that created it. Data in these sets will likely be sensitive and your security should reflect it. However, we need to determine rules for how the data should be modified, obfuscated, or deleted in order to make it consumable for broader audience, or what we might call “Silver” status. You need to create business rules to manage data (e.g. birthdays should be masked and social security numbers should be stored as only the last 4). Collibra is an example of a tool for this rules definition and management. It allows data rules to be set up based on logical business entities by business people rather than technologists.
  4. For those people who are familiar with governance concepts, you will recognize the difference between a policy and a control. A policy is like a speed limit sign along the highway. The control is the police officer that pulls you over if you are driving over that speed limit. Data works the same way. While Collibra establishes the policy, you need to create a method for enforcing that policy. To do this, you need to find the logical entities buried in the data (i.e. “oh look, I found a social security number!”). Examples of such products include Global IDs for scanning structured data sets with the operational systems and Waterline for scanning data inside of Hadoop.
  5. Once you have found the data that you want, you want to implement the rules. For this, there is an open source tool called Atlas that contains an orchestration capability called Falcon that helps implement the rules.
    1. Apache Atlas is a scalable and extensible set of core foundational governance services that enables enterprises to effectively and efficiently meet their compliance requirements within Hadoop and allows integration with the complete enterprise data ecosystem.
    2. Apache Falcon is a data governance engine that defines, schedules, and monitors data management policies. Falcon allows Hadoop administrators to centrally define their data pipelines, and then Falcon uses those definitions to auto-generate workflows in Apache Oozie
  6. Now that the data is loaded, you will want to enforce security through your LDAP capability or possibly through Kerberos. There are also tools like Blue Talon that simplify the ability to authorize, provision, protect, enforce and audit data security policies across your data lake.
  7. Finally, audit controls are critical. Cloudera introduced Navigator specifically to allow simple transparency to data history and lineage. Hortonworks will rely on Atlas to provide this capability.

Data that has gone through the above processes creates a view and accessibility of the data that can be made available to a wide set of users – both business analysts and data science teams.

Summary
When you build a house, the vast majority of the creative work is in the features and curbside appeal. That’s the fun part. But without the underlying plumbing, the house would quickly degrade into a money pit.

Consider the metaphor of a retail store: stocking the shelves vs. purchasing goods. When you go to the store, you don’t care about how the goods got there, but the rules for accessing the goods are everywhere. Cigarettes are behind the front desk. Pharmaceuticals must be dispensed with a prescription. Razor blades are under lock and key (for some strange reason). There are policies and enforcements on stocking the shelves so that the shopping experience is clear and easy.

To successfully operationalize the data lake, organizations need to address all of the plumbing requirements outlined in this blog that enable the business users and data science teams to have confidence in the wealth of data that the organization is amassing. The data lake plumbing processes may not be very glamorous, but without them, you’ll end up with a stinky data dump instead of a glorious data lake.

References:

  1. A “silver bullet” is a simple and seemingly magical solution to a complicated problem.
  2. While I mention several tools, this blog is not meant to be an endorsement of these tools nor is this intended to be a comprehensive list of such tools. However, many of these tools are the same tools that we use in our data lake implementations at EMC.

Data Lake Plumbers: Operationalizing the Data Lake
Bill Schmarzo

More Stories By William Schmarzo

Bill Schmarzo, author of “Big Data: Understanding How Data Powers Big Business”, is responsible for setting the strategy and defining the Big Data service line offerings and capabilities for the EMC Global Services organization. As part of Bill’s CTO charter, he is responsible for working with organizations to help them identify where and how to start their big data journeys. He’s written several white papers, avid blogger and is a frequent speaker on the use of Big Data and advanced analytics to power organization’s key business initiatives. He also teaches the “Big Data MBA” at the University of San Francisco School of Management.

Bill has nearly three decades of experience in data warehousing, BI and analytics. Bill authored EMC’s Vision Workshop methodology that links an organization’s strategic business initiatives with their supporting data and analytic requirements, and co-authored with Ralph Kimball a series of articles on analytic applications. Bill has served on The Data Warehouse Institute’s faculty as the head of the analytic applications curriculum.

Previously, Bill was the Vice President of Advertiser Analytics at Yahoo and the Vice President of Analytic Applications at Business Objects.

@ThingsExpo Stories
A strange thing is happening along the way to the Internet of Things, namely far too many devices to work with and manage. It has become clear that we'll need much higher efficiency user experiences that can allow us to more easily and scalably work with the thousands of devices that will soon be in each of our lives. Enter the conversational interface revolution, combining bots we can literally talk with, gesture to, and even direct with our thoughts, with embedded artificial intelligence, whic...
SYS-CON Events announced today that Cloudistics, an on-premises cloud computing company, has been named “Bronze Sponsor” of SYS-CON's 20th International Cloud Expo®, which will take place on June 6-8, 2017, at the Javits Center in New York City, NY. Cloudistics delivers a complete public cloud experience with composable on-premises infrastructures to medium and large enterprises. Its software-defined technology natively converges network, storage, compute, virtualization, and management into a ...
Every successful software product evolves from an idea to an enterprise system. Notably, the same way is passed by the product owner's company. In his session at 20th Cloud Expo, Oleg Lola, CEO of MobiDev, will provide a generalized overview of the evolution of a software product, the product owner, the needs that arise at various stages of this process, and the value brought by a software development partner to the product owner as a response to these needs.
New competitors, disruptive technologies, and growing expectations are pushing every business to both adopt and deliver new digital services. This ‘Digital Transformation’ demands rapid delivery and continuous iteration of new competitive services via multiple channels, which in turn demands new service delivery techniques – including DevOps. In this power panel at @DevOpsSummit 20th Cloud Expo, moderated by DevOps Conference Co-Chair Andi Mann, panelists will examine how DevOps helps to meet th...
Most technology leaders, contemporary and from the hardware era, are reshaping their businesses to do software in the hope of capturing value in IoT. Although IoT is relatively new in the market, it has already gone through many promotional terms such as IoE, IoX, SDX, Edge/Fog, Mist Compute, etc. Ultimately, irrespective of the name, it is about deriving value from independent software assets participating in an ecosystem as one comprehensive solution.
SYS-CON Events announced today that EARP will exhibit at SYS-CON's 20th International Cloud Expo®, which will take place on June 6-8, 2017, at the Javits Center in New York City, NY. "We are a software house, so we perfectly understand challenges that other software houses face in their projects. We can augment a team, that will work with the same standards and processes as our partners' internal teams. Our teams will deliver the same quality within the required time and budget just as our partn...
SYS-CON Events announced today that delaPlex will exhibit at SYS-CON's @ThingsExpo, which will take place on June 6-8, 2017, at the Javits Center in New York City, NY. delaPlex pioneered Software Development as a Service (SDaaS), which provides scalable resources to build, test, and deploy software. It’s a fast and more reliable way to develop a new product or expand your in-house team.
SYS-CON Events announced today that Tappest will exhibit MooseFS at SYS-CON's 20th International Cloud Expo®, which will take place on June 6-8, 2017, at the Javits Center in New York City, NY. MooseFS is a breakthrough concept in the storage industry. It allows you to secure stored data with either duplication or erasure coding using any server. The newest – 4.0 version of the software enables users to maintain the redundancy level with even 50% less hard drive space required. The software func...
SYS-CON Events announced today that A&I Solutions has been named “Bronze Sponsor” of SYS-CON's 20th International Cloud Expo®, which will take place on June 6-8, 2017, at the Javits Center in New York City, NY. Founded in 1999, A&I Solutions is a leading information technology (IT) software and services provider focusing on best-in-class enterprise solutions. By partnering with industry leaders in technology, A&I assures customers high performance levels across all IT environments including: mai...
With major technology companies and startups seriously embracing Cloud strategies, now is the perfect time to attend @CloudExpo | @ThingsExpo, June 6-8, 2017, at the Javits Center in New York City, NY and October 31 - November 2, 2017, Santa Clara Convention Center, CA. Learn what is going on, contribute to the discussions, and ensure that your enterprise is on the right path to Digital Transformation.
In his keynote at @ThingsExpo, Chris Matthieu, Director of IoT Engineering at Citrix and co-founder and CTO of Octoblu, focused on building an IoT platform and company. He provided a behind-the-scenes look at Octoblu’s platform, business, and pivots along the way (including the Citrix acquisition of Octoblu).
SYS-CON Events announced today that Systena America will exhibit at SYS-CON's 20th International Cloud Expo®, which will take place on June 6-8, 2017, at the Javits Center in New York City, NY. Systena Group has been in business for various software development and verification in Japan, US, ASEAN, and China by utilizing the knowledge we gained from all types of device development for various industries including smartphones (Android/iOS), wireless communication, security technology and IoT serv...
SYS-CON Events announced today that Outscale will exhibit at SYS-CON's 20th International Cloud Expo®, which will take place on June 6-8, 2017, at the Javits Center in New York City, NY. Outscale's technology makes an automated and adaptable Cloud available to businesses, supporting them in the most complex IT projects while controlling their operational aspects. You boost your IT infrastructure's reactivity, with request responses that only take a few seconds.
DevOps at Cloud Expo – being held October 31 - November 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA – announces that its Call for Papers is open. Born out of proven success in agile development, cloud computing, and process automation, DevOps is a macro trend you cannot afford to miss. From showcase success stories from early adopters and web-scale businesses, DevOps is expanding to organizations of all sizes, including the world's largest enterprises – and delivering real r...
SYS-CON Events announced today that CollabNet, a global leader in enterprise software development, release automation and DevOps solutions, will be a Bronze Sponsor of SYS-CON's 20th International Cloud Expo®, taking place from June 6-8, 2017, at the Javits Center in New York City, NY. CollabNet offers a broad range of solutions with the mission of helping modern organizations deliver quality software at speed. The company’s latest innovation, the DevOps Lifecycle Manager (DLM), supports Value S...
SYS-CON Events announced today that Enzu will exhibit at SYS-CON's 20th International Cloud Expo®, which will take place on June 6-8, 2017, at the Javits Center in New York City, NY, and the 21st International Cloud Expo®, which will take place October 31-November 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. Enzu’s mission is to be the leading provider of enterprise cloud solutions worldwide. Enzu enables online businesses to use its IT infrastructure to their competitive ad...
SYS-CON Events announced today that Peak 10, Inc., a national IT infrastructure and cloud services provider, will exhibit at SYS-CON's 20th International Cloud Expo®, which will take place on June 6-8, 2017, at the Javits Center in New York City, NY. Peak 10 provides reliable, tailored data center and network services, cloud and managed services. Its solutions are designed to scale and adapt to customers’ changing business needs, enabling them to lower costs, improve performance and focus intern...
Everywhere we turn in our industry we can find strong opinions about the direction, type and nature of cloud’s impact on computing and business. Another word that is used in every context in our industry is “hybrid.” In his session at 20th Cloud Expo, Alvaro Gonzalez, Director of Technical, Partner and Field Marketing at Peak 10, will use a combination of a few conceptual props and some research recently commissioned by Peak 10 to offer a real-world consideration of how the various categories of...
In order to meet the rapidly changing demands of today’s customers, companies are continually forced to redefine their business strategies in order to meet these needs, stay relevant and continue to see profitable growth. IoT deployment and development is integral in this transformation, and today businesses are increasingly seeing the value of investing their resources into IoT deployments. These technologies are able increase ROI through projects such as connecting supply chains or enabling sm...
SYS-CON Events announced today that SoftLayer, an IBM Company, has been named “Gold Sponsor” of SYS-CON's 18th Cloud Expo, which will take place on June 7-9, 2016, at the Javits Center in New York, New York. SoftLayer, an IBM Company, provides cloud infrastructure as a service from a growing number of data centers and network points of presence around the world. SoftLayer’s customers range from Web startups to global enterprises.