Jake Watson, Author at Oakland Thu, 25 Sep 2025 07:55:12 +0000 en-GB hourly 1 https://wordpress.org/?v=6.9.4 https://weareoakland.com/wp-content/uploads/2024/01/cropped-oakland-favicon-150x150.jpg Jake Watson, Author at Oakland 32 32 Should you Adopt Microsoft Fabric as your Data Platform? https://weareoakland.com/blog/adopt-microsoft-fabric-as-your-data-platform/ Tue, 07 Jan 2025 14:29:08 +0000 https://weareoakland.com/?p=9239 As Microsoft Partners; here at Oakland we’re being asked by more and more clients if Microsoft Fabric could help them to drive better more profitable outcomes from their data? Source: https://www.microsoft.com/en-us/microsoft-fabric Microsoft Fabric has been out on general release for a while now and is increasingly seen as a serious alternative to a bespoke Data...

The post Should you Adopt Microsoft Fabric as your Data Platform? appeared first on Oakland.

]]>
As Microsoft Partners; here at Oakland we’re being asked by more and more clients if Microsoft Fabric could help them to drive better more profitable outcomes from their data?

Source: https://www.microsoft.com/en-us/microsoft-fabric

Microsoft Fabric has been out on general release for a while now and is increasingly seen as a serious alternative to a bespoke Data Platform. Promising the capabilities of a one stop shop for all your data needs, from simple dashboards to advanced data science workloads. But is it right for you? Even if you currently use the Microsoft data stack (Synapse, on-premises SQL Server) the answer isn’t an automatic yes, as we’ll explain below…

What the Problems Does Microsoft Fabric Solves?

For existing Microsoft customers, Fabric is certainly being touted as a no-brainer. Integrating with Microsoft 365, enabling collaboration through Teams, Copilot and SharePoint users can quickly access insights directly using familiar tools. 

Managing multiple tools for data visualisation, governance, and collaboration can be cumbersome and costly therefore bringing them all under the Microsoft umbrella can be seen as easier and more cost effective.

The popular Power BI integration can provide visualisations and reporting tools directly tied into the platform.

And let’s not forget Data Governance – Fabric promises robust data governance, ensuring compliance, security, and data lineage across all processes.

  • Built-in Security: Unified security measures like role-based access and encryption.
  • Unlike several alternative data platforms, there is no extra cost or much additional configuration for Microsoft Entra/Office 365 Single Sign-On (SSO).
  • Data Lineage: Comprehensive tracking of data transformations and movement.
  • Purview Integration: Facilitates data cataloguing and compliance tracking.

Fabric uses OneLake, a universal data lake that centralises data storage and enables cross-platform access without creating silos. This streamlines data integration from multiple sources, including on-premises, cloud, and SaaS systems.

Many organisations we see struggle with siloed data across on-premises, cloud, and SaaS platforms, making integration complex and time-consuming. A traditional data platform can be the answer but often by the time the platform is created the technology is out of date or the people and processes needed to use the platform to its full potential are disengaged and doing their own thing.

Direct Lake Implementation, while having a few limitations (largely it’s not great for reasonably complex Power BI dashboards at present) does offer near real-time loading of large data sets with little to no configuration, which can be a major productivity boost. How often have your data teams been left twiddling their thumbs waiting minutes or even hours for your Power BI datasets to refresh because of changes made to your source data with direct import?

And as 2025 is the year of Agentic AI we can’t fail to mention the challenge of incorporating AI into existing workflows. Of course, Microsoft Fabric has a solution for that with its AI Co-pilot integration, while only currently available at F64 capacity / price tier (£8k a month) this gives us a tantalising glimpse into major productivity boosts for your engineers and analysts. Sound good so far? These are also improvements over Microsoft’s other data platform, Azure Synapse Analytics, but there is also improved Spark performance, more integrated real-time analytics, Lakehouse shortcuts for easier loading of data in the Data Lake, plus many more features and by the time this blog is published no doubt Microsoft will have added more new features to the mix.

How Hard is it to Implement Fabric?

Getting started super easy: deploying a capacity and workspace only requires a few clicks.

Although we’ve noticed implementing a typical development, test, and production software lifecycle in Fabric can be tricky as git integration and/or pipeline deployments are either in preview or straight up missing for some Fabric features.

Source: https://learn.microsoft.com/en-us/fabric/cicd/git-integration/intro-to-git-integration?tabs=azure-devops#supported-items

You also must be careful how you implement Fabric; it has many options! For example, there are three different ways of doing data transformations:

  • Python Notebooks
  • SQL Queries
  • Dataflow Gen 2 (low-/no-code)

Each has their own use cases, personas, pros and cons, so we recommend you understand each of them, including trying them out before committing to one (or more). This goes for many other aspects of Fabric (ingesting data, CI/CD process, data science workflow, etc.).

Another thing to note is that if you want private endpoints or on-premises data, they are more work to implement and can be tricky to do, especially if you’ve not implemented them before.

What does Fabric Cost?

Unlike Synapse, Databricks and Snowflake which offer pay-as-you-use/pay-as-you-go with auto-shutdown for compute costs, Fabric compute capacity charges are like a cloud Virtual Machine (VM) or a Netflix subscription: it’s always on and racking up costs unless you pause it, which turns off all compute and most of its features.

Just like a cloud VM, they are quick and easy to build, start, pause, and destroy via the Azure portal and can build more than one of them in an Azure tenant, so we find them great for testing out ideas.

Source: https://learn.microsoft.com/en-us/azure/architecture/analytics/architecture/fabric-deployment-patterns#pattern-3-multiple-workspaces-backed-by-separate-capacities

This does mean, however, production Fabric capacities are often left on 24/7, which means more simple and predictable cost forecasts but potentially not as efficient if you’re not using all or most of your capacity 24/7 as serverless pay-as-you-use, especially as serverless Databricks and Snowflake are generally cheaper per hour than their non-serverless workloads.

Common Data Platform compute costs throughout a typical (Mon-Fri) working day.

To be fair, if you’re running Fabric 24/7, you can get 41% discount for a year’s reservation and you could scale & pause compute on a schedule with Fabric’s REST API, especially for development workloads. Though doing this starts to impact Fabric’s main selling point – it’s easy to run and maintain.

UK South Capacity Pricing. Source: https://azure.microsoft.com/en-us/pricing/details/microsoft-fabric/

A single compute capacity can also cause problems as you scale in users and workloads, as workload management is a bit limited in Fabric right now, and your capacity compute powers almost everything. This can lead to trying to figure out who or what has eaten all the compute. This can also lead to further costing inefficiencies as you scale up pricing tier /capacity or cloning workspaces to another capacity avoid having your work blocked by someone else at busy periods.

There are other potential Fabric costs to be aware of:

  • OneLake storage costs, which are similar or slightly higher (-0 to 20%) than Azure Data Lake costs.
  • You’ll still need Power BI Pro licences for your Power BI developers.
  • Gateway costs if you require a vNet/on-premises gateway.
  • Private Endpoint costs, if required.

The above costs are often small compared to capacity costs, with capacity costs often being approx. 70% to 95% of your total run costs for Fabric depending on the options you choose.

Also note, many preview features are locked behind F64 capacity / pricing tier: we’ve seen some clients opt for this capacity to get the additional features, even though they don’t need the compute.

Who is Fabric Suitable For?

Existing Fabric/Power BI users who want to expand their data capability without going through the extra hassle of adopting another tool.

Migrating from low-/no-code focused products then Atleryx and Talend are alternative options, especially if you’re looking for a data platform with more Azure integration.

Migrating from an on-premises or cloud SQL Server will give you more data engineering, science and governance capability compared to on-premises tooling like SSIS.

You might also think existing Synapse users are great for Fabric since they are made by the same company, and Fabric is in many ways an evolution of the Synapse features previously mentioned. This is very true: a migration to Fabric will be easier than to any other product. Though beware, as mentioned above the cost model is different in Fabric, and there are some features in Synapse not or heavily changed in Fabric* that could block a migration or require significant rework (for example a lack of parameters in pipeline connections).

We don’t believe existing or target (large enterprises) Databricks or Snowflake customers will be adopting Fabric immediately: while the foundations are there, I think there are too many essential features in preview or missing, with examples such as git integration and workload management, which I mentioned above.  

But and there is always a but the only caveat we’d add is if you have a business need for using Direct Lake, we can see some large companies using Fabric with Databricks/Snowflake in tandem – Microsoft even makes this easy by having great integration with Databricks and Snowflake.

Source: https://learn.microsoft.com/en-us/fabric/onelake/onelake-overview

*As of January 2025, we suspect this sentence will need changing in the coming months as Microsoft is working hard for feature parity with Synapse.

Summary

While Fabric is a step forward from Synapse, there are differences you’ll have to think about before migrating, and it doesn’t make Azure Databricks redundant either, especially for large, complex enterprises.

And while we’ve found it a great experience to get started with Fabric, there can be a few challenges and decisions to be made before reaching production.

If you would like a deeper dive into comparison of Fabric against its competitors and advice on how best to implement Fabric, feel free to reach out or download our Data Platform guide.

Interested in unifying your data estate with Microsoft Fabric? Join our workshop to discover how Microsoft Fabric can streamline your data integration and management. We’ll explore its features and assess whether it’s the ideal solution for your organisation’s needs.

The post Should you Adopt Microsoft Fabric as your Data Platform? appeared first on Oakland.

]]>
Is Microsoft Purview The Answer to Modern Data Governance? https://weareoakland.com/blog/is-microsoft-purview-the-answer-to-modern-data-governance/ Mon, 30 Sep 2024 09:30:46 +0000 https://weareoakland.com/?p=9070 In today’s data-driven world, effective governance is crucial for managing vast volumes of information while ensuring security and compliance. Microsoft Purview promises a robust solution to modern data governance challenges, offering seamless integration with Microsoft’s ecosystem. As businesses increasingly rely on data to drive decision-making, tools like Purview help streamline data cataloguing, classification, and protection. ...

The post Is Microsoft Purview The Answer to Modern Data Governance? appeared first on Oakland.

]]>
In today’s data-driven world, effective governance is crucial for managing vast volumes of information while ensuring security and compliance. Microsoft Purview promises a robust solution to modern data governance challenges, offering seamless integration with Microsoft’s ecosystem. As businesses increasingly rely on data to drive decision-making, tools like Purview help streamline data cataloguing, classification, and protection. 

But is it the right fit for your organisation? In this blog, we explore Purview’s features, deployment, and pricing, helping you determine if it can meet your data governance needs while ensuring compliance and scalability.

How Do You Govern Your Data?

Most organisations are exploding with data that has been collected, transformed, and reported on with the business requirement to improve decision making. However, this huge increase in the volume of data has come with a lack of accurate tracking which all too often hampers the actionable insights that the business stakeholders demand. As organisations become more data-driven, Oakland has seen a growth in 4 particular pains which have been increasing for the last few years: 

  • How can we audit all this data to protect against data leaks and unexpected data loss which is especially crucial with the regulatory requirements of many organisations? 
  • How can data users discover data and derive business value in an environment that changes constantly? 
  • How can data consumers understand what the data collected means and turn this into business value? 
  • How can we show the current data quality of key datasets? 

Plus, in a modern business environment, you may need on-premises and multi-cloud data governance solution which also easily integrates with your office 365 workloads.

Data Governance tooling can help mitigate these problems, and help with data management however, these tools are often complex to integrate to your entire data estate due to requiring: 

  • The ability to scan a large variety of data sources. 
  • A highly customised user interface. 
  • A powerful search engine to find data assets by many different types of metadata attributes. 
  • Technical experience to setup a Catalog, and Data Stewards with experience on maintaining a Catalog. 

These are just a few of the main requirements that create a software marketplace full of products that are often expensive and hard to implement and maintain which is why these are often not business friendly and/or very expensive.  

These products also need to ingest large amounts of sensitive data to meet user requirements, ironically creating a data governance concern in itself! 

Microsoft Purview aims to ease the pain of data governance by being feature-rich, easy to deploy, maintain and secure. But is it worth the cost, and can it compete with bespoke data governance companies that have a head start measured in years or even decades? 

What Are the Features of Microsoft Purview? 

Microsoft Purview has recently gone through a major new update which adds lots of features, so if you’ve dismissed Microsoft Purview before, we recommend looking again. 

  • It’s connectors are very Microsoft-focused but cover most of its ecosystem: Azure, SQL Server, Power BI, and Office 365. If you’ve already bought heavily into Microsoft, you can scan most or all your data assets automatically. 
  • It focuses less on connectors made by other companies but still covers many popular data products like SAP, Salesforce, Oracle, GCP Big Query, AWS S3, and Snowflake.  
  • You can create and streamline business domains like sales, marketing, HR and supply chain, to help move data governance closer to the business  
  • Data classification becomes more business friendly as you can classify data with 200+ pre-built classifications, as well as custom classifications. 
  • The pre-built data governance compliance reports enable you to quickly check insights such as the percentage of data that has a data owner and the percentage of new data assets in the last month. 
  • API and Python SDK enable you to create custom data sources where connectors don’t exist or mass updating existing scanned data assets. 
  • Data quality is always a key issue Microsoft Purview has pre-defined and custom data quality tests. 
  • Data Sharing allows users to give other users read-only data lake data access without having to copy data. 
  • Data Governance can integrate with Master Data Management tooling like Profisee 
  • AI powered with lots of integration with Microsoft Copilot to generate data documentation and data quality tests. 

How Do You Deploy Microsoft Purview? 

  • Oakland has designed and built many data platforms, and we highly value any product that can be deployed quickly, has low maintenance, and will meet stringent client IT & security requirements. We believe Microsoft Purview is stronger than most data governance products when it comes to data protection. 
  • It is as easy to deploy and maintain in Azure as any SaaS data governance product but also offers a choice – 20 plus regions to deploy into, including the UK. 
  • Microsoft Purview can also keep all traffic in and out of its server on its private network using Private Endpoints, never touching the public internet, offering an extra layer of data security when creating a Data Catalog. 
  • Scan Azure data via Managed Identity authentication which offers high-security data connections without worrying about managing passwords. 
  • It can connect directly to scan on-premises and other public cloud data assets (for example, AWS and GCP), though it requires some technical knowledge to setup the networking. 

How Much Does Microsoft Purview Cost? 

Automated data governance tooling is expensive, with costs starting in the thousands of pounds for most products. Microsoft Purview arguably starts at a lower base: we’ve found it starts at about £250 per month. However, you will also be charged on top of the base cost for scanning data, which goes up the more data consumed. 

Due to the pricing being highly variable in Microsoft Purview we recommend building a proof of concept to road-test Microsoft Purview for a month or so to accurately measure costs. 

What Are The Alternatives to Microsoft Purview?

Note this isn’t a comprehensive list and is a quickly evolving space with new exciting start-ups, and apps entering all the time, but we hope it will help you make an informed decision. 

  • Build your own: Building your own data governance tool offers complete customisation to your business’s unique needs, allowing you to tailor features and integrate seamlessly with existing systems. It gives you full control over security and privacy, ensuring sensitive data stays in-house, and avoids costly vendor licensing fees, providing better long-term ROI. You also avoid vendor lock-in, giving you the flexibility to evolve your tool as your governance requirements change.

While developing a custom tool requires a significant upfront investment, it fosters internal expertise and provides faster iteration when compliance needs or data regulations shift. By owning the solution, your organisation gains greater agility and control, avoiding reliance on third-party support or updates.

  • Excel: low cost, low maintenance if data structures don’t update regularly, doesn’t require any specialist skills to build. While we suspect this is the most common type of data catalog used, we feel nervous about doing a data catalog in a data tool infamous for having poor data governance. It does not scale and requires lots of manual effort for any major changes to the organisation, creating data lineages and classification of sensitive data. 
  • Automate your own solution by extracting schemas of databases and files. This is a nice quick way of generating a data catalog with low maintenance and little extra costs. You can also build a dashboard on top of the business intelligence (BI) platform of your choice. It requires minimum effort if the number of data assets is small. Although adding features like data lineage and classifying data sensitivity will require a reasonable amount of engineering effort, which makes buying off the shelf products more appealing.
    • You can combine this with a SharePoint Site to collect business information like Business Domains to make sure it is not just an IT exercise. 
  • Databricks Unity Catalog – ideal for Databricks heavy data platforms, as it is a free extra. Though it will only scan what Databricks can scan. You can integrate with other data governance products, including Microsoft Purview, and update schema as they update in real time.  
  • Mature products like Informatica and Talend. These tend to charge by the user and are more commonly found on-premise (though they can be configured and maintained in the cloud on Virtual Machines). They will likely cost the most; sometimes, this is significant, but these are feature rich, well-trusted and reliable.  
  • New(er) products like Atlan and Immuta often focus on providing data governance to more recent cloud data tooling like Databricks and Snowflake but also often focus on making deployments into the cloud more accessible by offering deployments via Docker or Kubernetes.
    • Immuta also provides a single pane of glass for fine-grain data access across many popular data products that allows data access controls at a column and row level. 
  • Open Source software like Datahub and Amundsen, both built by large tech companies (LinkedIn and Lfyt respectively). These are the go solutions if your organisation has the technical capacity to build and maintain complex workflows. They offer a lot of customisations and the possibility of no licence costs, so they can be much cheaper at scale and be more custom tailored to fit an organisation’s data governance needs. 

If you would like to see more tooling options and deep dive into how to select the right data governance tool to streamline your governance practices check out our data governance tooling guide.

Why Should You Choose Microsoft Purview?

In an increasingly busy data governance market, Azure Purview is a serious option to consider especially considering the new update which fills in some major gaps in its features. 

If you are looking for a data governance product that is easy to deploy, secure, catalogue, and classify data assets, and provides some customisation through APIs and user interface at a competitive cost, then we think Microsoft Purview is a good contender. 

However, and there is always a but, we do want to end on a cautionary note: we have found Microsoft Purview or indeed any other data catalog implementation fails more often because of either:  

a) Lack of data governance processes and people.  

b) Not knowing what business problems, you are exactly trying to solve. 

Rather than choosing the wrong tool. This is because data catalog requires constant maintenance to keep up with the evolving nature of any organisation, so needs to show a high level of return of investment and have the right processes and people to maintain it efficiently.  

Also remember to make sure that you select a tool based on the problems you have and how the tool can help you solve them and keep reviewing how it does can could add value.  If you would like to know more about how we help clients with data governance while achieving a quick return on investment, download our data governance guide.   

The post Is Microsoft Purview The Answer to Modern Data Governance? appeared first on Oakland.

]]>
Should you use Data Lakehouse instead of a Data Warehouse and / or Data Lake?  https://weareoakland.com/blog/should-you-use-data-lakehouse-instead-of-a-data-warehouse-and-or-data-lake/ https://weareoakland.com/blog/should-you-use-data-lakehouse-instead-of-a-data-warehouse-and-or-data-lake/#respond Fri, 08 Mar 2024 14:39:42 +0000 https://www.theoaklandgroup.co.uk/?p=7265 Should you use Data Lakehouse instead of a Data Warehouse and / or Data Lake?  Intro When using your Data Platform to improve your Business Intelligence with useful dashboards and reports, you’ll more than likely want to use a Data Warehouse. Add on your data science builds, and storing your raw data cheaply, plus adding...

The post Should you use Data Lakehouse instead of a Data Warehouse and / or Data Lake?  appeared first on Oakland.

]]>
Should you use Data Lakehouse instead of a Data Warehouse and / or Data Lake? 

Intro

When using your Data Platform to improve your Business Intelligence with useful dashboards and reports, you’ll more than likely want to use a Data Warehouse. Add on your data science builds, and storing your raw data cheaply, plus adding a Data Lake just for good measure, and the costs soon start to add up. Running both in tandem on a Data Platform can have serious costs and maintenance associated.

So, can you have the best of both worlds with the Data Lakehouse? And what is the best Lakehouse to use

Before we answer those questions, we must ask “What is a Data Warehouse, Data Lake and a Data Lakehouse?”

What is a Data Warehouse, Data Lake and a Data Lakehouse?

Data Warehouse is a data architecture that has been around since the 90s and is still relevant today. It is a means to store tabular data so it can be easily used by business intelligence applications such as Tableau or Power BI, web applications, and even other data warehouses. The three most common Data Warehouse architectures are Kimball Star Schema, Data Vault and One Big Table.

The name is also confusingly used to identify a type of Database, such as AWS Redshift, Azure Synapse and Snowflake, which specialise in storing and querying large amounts of data.

Data Warehouses have their issues; they can be more expensive than a Data Lake when processing large amounts of data, and work best when data is of reasonable quality and in a tabular structure.

Architecture of a simple Data Platform using just a Data Warehouse. 

So, along came the Data Lake to help ease these common pain points:

  • Data Scientists needing to be able to process large amounts of raw data of dubious quality. 
  • Increasing requirements for storage of non-tabular data sources. 
  • The need for data storage that is more flexible in structure and schema. 
  • The need to store data that might be needed at later date, for example for auditing, but have a low setup and maintenance cost (little or no ETL process needed compared to a Database). 

Data Lake is just a distributed file system at its heart, usually hosted in the cloud in AWS S3 or Azure Data Lake, with large files split by a key, so you can save on processing costs by loading the partitions you need.

Data Lakes also generally have more flexibility in that it can store an unlimited number of file formats and offer a common interface to its storage that allows you to use many compute engines. This often called separating storage from compute, which has become so popular that many Data Warehouses offer this too now. Data Lakes can also easily store non-tabular data (images, videos and music) that Data Warehouses cannot without some pre-processing.

However, without Delta Lake it cannot easily or efficiently do row level updates and inserts, nor connect easily to business intelligence applications, that a Data Warehouse or Database can do.

Architecture of an example Data Platform using both a Data Lake and Data Warehouse.

What is a Data Lakehouse?
A Data Lakehouse is an open data management architecture that combines the flexibility, cost-efficiency, and scale of Data Lakes with the data management and ACID transactions of Data Warehouses, enabling business intelligence (BI) and machine learning (ML) on all data.  

What is Databricks Lakehouse? 
Until a few years ago, Databricks was mainly designed as an easy way to run Spark, a distributed data processing library for large scale Data Engineering and Data Science. It worked mainly in tandem with a Data Lake, with similar advantages and drawbacks. 

In 2019 Databricks released Delta Lake, a file format with attributes only found previously in Databases and Data Warehouses as mentioned above. Combined with Spark to process and transform a wide variety of data, this gave birth to the Data Lakehouse. 

Today, Databricks has a fully featured SQL Data Warehouse, enterprise security, data governance with Unity Catalogmany data connectors, as well as the ability to output data to Power BI and Tableau, so it can meet all common data use cases. 

Architecture of an example Databricks “Lakehouse” using Spark as the processing engine and Delta Lake as storage. 

For those looking at building a Data Mesh, Databricks has federated query in preview, though Delta Lake also has connectors for TrinoStarburst and Dremio so you can join up many Data Lakes across your organisation: 

Architecture of many Lakehouse Data Products in a Data Mesh – the query layer and governance layer will have access to all Data Products, limited by access permissions.  

Will I still need a Data Warehouse? 

Maybe, but note it may take some time for a data team used to Databases/Data Warehouses and SQL to convert to Data Lakehouse. Here at Oakland we feel it is still easier to set up and optimise Cloud Native Warehouses like Snowflake and Google Big Query, than Databricks, as there are fewer moving parts.  

These maintenance costs can far outweigh the benefits of the Lakehouse, generally at smaller scales and data complexity. 

Also, while we’ve seen first-hand that Lakehouse can be the cheaper and more performant option than a Data Warehouse, this hasn’t been the case 100% of the time and you should do your own testing, as performance and cost heavily depends on the data you use and the environment you operate in. 

Can I build a Lakehouse somewhere other than Databricks? 

Yes, Delta Lake is open source and can be used in many different data compute products which are listed below. However, Databricks has built in special optimisations just for Databricks and a robust user interface to manage the Lakehouse. So, it is likely running Delta Lake will be slower and could be harder to maintain elsewhere. 

Example the Databricks user interface for datasets showing the schema and a sample of the dataset. 

Also note that Databricks is a general compute engine rather than a database or programming interface: it can run SQL, Pandas, Ray, Spark, most of the popular data science libraries, do graph analytics, geospatial, IoT, near-real time streaming and import almost any Python, Java, R or Scala library. Databricks’ main benefit to us is its extreme versatility, potentially reducing costs by not having to maintain separate business intelligence and data science data processing applications.    

Also, Databricks is in strong position to customise Large Learning Models (LLMs) like ChatGPT, with its general compute and strong MLflow integration, so you can pick the best open-source AI models and tune it with your organisational data in a highly efficient way using MLOps.  

However, if you’re already using one of the Lakehouse alternatives listed below, it may not be worth adding Databricks to your Data Platform. 

The alternatives to Databricks Lakehouse are:  

  • Starburst, like Databricks, is a cloud neutral and cloud native compute engine with a full suite of enterprise options and data connectors. It has Delta Lake and Iceberg connectors that can be fully controlled with a SQL API.  
  • Azure Synapse has the option to use its own Spark Engine, can import Java and Python libraries, and has Delta Lake Integration too. Has excellent integration with rest of Azure.  
  • AWS Glue allows you to use Delta Lake in S3. Has excellent integration with rest of AWS. 

Some may say Pandas or DuckDB can be a Data Lakehouse, though from our research in May 2023 they cannot do transactions or merges on a Data Lake file (Delta Lake, Iceberg, etc.) so have been excluded from the above – they still have their own use cases though.  

Summary

In short, like with other data products and architectures, the answer is it depends on the makeup of your data team, security, the size and structure of your data, and how the data is used among many other factors.  

If you are consuming a lot of data in your data platform, struggling to manage both a Data Lake and Data Warehouse at the same time, or trying to figure out how to use advanced analytics like Machine Learning with your data, Data Lakehouse is in our opinion a convincing proposition.   

We also find ourselves recommending Databricks more often than the alternatives as it offers the most complete Lakehouse solution, though competitors are quickly catching up and offering a near as good as experience as Databricks, so the choice isn’t as easy to make as it was in 2021 when we first wrote this article. 

The post Should you use Data Lakehouse instead of a Data Warehouse and / or Data Lake?  appeared first on Oakland.

]]>
https://weareoakland.com/blog/should-you-use-data-lakehouse-instead-of-a-data-warehouse-and-or-data-lake/feed/ 0
Should you build or buy your data platform? https://weareoakland.com/blog/should-you-build-or-buy-your-data-platform/ https://weareoakland.com/blog/should-you-build-or-buy-your-data-platform/#respond Wed, 08 Nov 2023 13:36:50 +0000 https://www.theoaklandgroup.co.uk/?p=7790 “Build Versus Buy” is More a Scale of Options than a Binary Decision As consultants, we’ve been involved in helping our clients to decide which is the best software to invest in (or not invest in some cases). This is quite a responsibility as we are putting our reputation in the hands of a vendor,...

The post Should you build or buy your data platform? appeared first on Oakland.

]]>
“Build Versus Buy” is More a Scale of Options than a Binary Decision

As consultants, we’ve been involved in helping our clients to decide which is the best software to invest in (or not invest in some cases). This is quite a responsibility as we are putting our reputation in the hands of a vendor, and we know that tech vendors are amazing at making claims. We see it in every trade show we attend. Therefore, it can be impossible to verify all the features of every product.

From our experience, this can be difficult to get right. So here are a few lessons in buying software (or not) to avoid making costly mistakes.

Limiting Your Potential Options is Sensible but Dangerous

One sales tactic is to present two options. This makes sense in some ways: management is time constrained and doesn’t want to assess all 100+ options that are possible, but this can lock you into a limited way of thinking, especially if you only assess two extreme “buy“ and “build” options.

Diagram to show investments in software for a data platform on a scale

Scale of Product Types from Build to Buy. By Jake Watson.

So you want to look at a variety of options, and for most organisations, look at options across the scale. We find a managed Platform as a Service (PaaS) in the middle of the scale, fits well for most but not all!

Think Beyond a Product

You may also want to consider a mix of both build and buy components; this can get you closer to fulfilling a set of complex requirements in a reasonable budget and timeframe.

One of our favourite examples is Fivetran, a vendor that offers a super quick and easy way of getting source data into your Lakehouse or Warehouse. However, it can cost a lot at scale, so you may end up doing the math and finding that building your own ingestion is more cost-effective for some, but not all, sources. Or try open source (AirbyteMeltano).

And yes, maintaining two or more systems is generally worse than one, but that should be a guideline, not a rule, as sometimes, you are the exception.

There are Too Many Choices Though!

You can go too far the other way and spend months going through dozens of options out of 1400+ products that exist in data and AI.

One example of the variety of options available to you is real time data processing, where you can choose from least to most managed:

*For those who haven’t come across IaaS, PaaS and SaaS we recommend this article.

We’ve probably missed a few more options, and this doesn’t even consider rivals to Kafka (Red Panda, Pulsar) We understand that this can feel overwhelming, so the sensible approach is to take a funnel approach and briefly compare lot of options to quickly to exclude and then narrow down to a few options for a deeper comparison.

Three of the biggest reasons we’ve come across for product exclusion are:

  • Too expensive for the budget
  • Bad fit for the team that will build and maintain the software (say, for example, knowledge of needing to use Java, in a team that only knows Python and SQL).
  • Doesn’t meet security requirements

So start here before comparing more technical requirements that will require more work to compare and make a list of must-haves in MOSCOW) before comparing nice-to-haves requirements (must haves in MOSCOW) before comparing nice-to-haves must have’s in MOSCOW) before comparing nice-to-haves to save on extensive comparisons.

Where to Start

Starting a product comparison can be the hardest part if you have no experience in the domain. There are a good few places to start: hiring expertise if you don’t have the time. Oakland is often called in to assess a client’s options. As we are tech-agnostic and have experience working with many different clients, we know what works and what doesn’t! (shameless plug for us), read blogs and books by subject matter experts (check out the Data Platform Journal another shameless plug) or use social media for advice.

Try to avoid straight-up sales pitches and salespeople early on; you are trying at this stage to gather information to make a better-informed decision later on, not make a decision right now. Most have useful guides and content which you can use to inform your decision-making.

Modern Data Stack is also a great website for finding a range of data solutions, though skewed towards smaller start-ups. Gartner is another option if you have a license, but it skews towards established buy solutions.

What to Watch Out For When Buying Software

Avoid bias towards slick sales staff and fancy user interfaces: vendors can hype up their products with years of sales experience, whereas engineers have little to no sales experience to hype up their build option.

Also, how much does a good User Interface (UI) for your users matter? For engineers, often not much, non-technical users, a lot.

Buyers may not realise that even heavily managed solutions require some configuration and support, not as much support as a build option, but this aspect is often forgotten about.

People tend to also forget about the 10% trap, where on average a complex solution requires 10% to 20% of its requirements to be met by highly customisable software.

We’ve seen a number of times people look at and/or test the simple use cases and forget the more difficult requirements they have, which inevitably means that the solution comes unstuck and can’t deliver on all of the requirements. Test your potential solutions against your most challenging must-have requirements.

Beware of your own Engineer’s Biases Towards Build

So far, this may sound a little biased towards “build“ solutions, so we’ll rein it in here.

Most experienced engineers would back themselves to build solutions that can also be bought if given the time and money, but just because you can doesn’t mean you should. Often, projects take longer and throw up difficulties along the way.

So, if the build and buy solutions cost roughly the same and both meet requirements, we would er on the side of caution and go for the buy option.

Another classic trap engineers fall into is pitching a solution that is technically better but doesn’t improve the organisation: making a pipeline run faster with fancy new tech doesn’t always make the organisation’s profits larger.

If you’re not an engineer, then you may be more likely to be biased towards a “buy“ option in our experience, though again, there are exceptions.

Summary

So, in summary: avoid extremes, gather information before engaging sales teams, compare using the most difficult and important requirements first, be creative and check your biases towards either buying or building.

Author: Jake Watson

The post Should you build or buy your data platform? appeared first on Oakland.

]]>
https://weareoakland.com/blog/should-you-build-or-buy-your-data-platform/feed/ 0
Why Invest in Data Quality? https://weareoakland.com/blog/why-invest-in-data-quality/ https://weareoakland.com/blog/why-invest-in-data-quality/#respond Wed, 23 Aug 2023 09:31:22 +0000 https://www.theoaklandgroup.co.uk/?p=7551 This can seem like a rhetorical question: you should always invest in Data Quality! But we are still not investing enough: surveys show Data Quality issues are increasing in most organisations and on average, take up 34% of a Data Engineers time instead of them creating value by adding new features. This increases to 50% in large...

The post Why Invest in Data Quality? appeared first on Oakland.

]]>
This can seem like a rhetorical question: you should always invest in Data Quality! But we are still not investing enough: surveys show Data Quality issues are increasing in most organisations and on average, take up 34% of a Data Engineers time instead of them creating value by adding new features. This increases to 50% in large Data Platforms.

All these Data Quality issues add up, with bad Data Quality costing organisations on average $15mil a year.

Data Quality investment is also an investment in high-quality AI and ML, as you’ll likely get more accurate AI and ML results from improving Data Quality than changing your AI model and code.

Having Data Quality checks in place helps reduce “data downtime” for outages and fixes, which subsequently increases the overall reliability of the Data Platform. Highly reliable data leads to more trust in data and better-informed decision-making.

Better decision-making should increase profitability, productivity, and confidence of the whole organisation, which in turn usually leads to more investment in data and, as a result, going back to the start: Increasing the quality of data again.

All this creates a “Virtuous Cycle” of Data Quality, constantly improving your organisation:

If Data Quality decreases the opposite happens with a negative cycle.

Better Data Quality testing should also reduce the blast radius of the issue to a few Data Engineers rather than hundreds or thousands of users as more issues are being found earlier:

The fewer users impacted, the smaller the cost caused by the issue, which should again pay back any investment in Data Quality in large multiples.

Do I need a Data Quality Framework?

I know that setting out to build a framework for anything requires time, and you’ll have many competing concerns, so we understand if you feel reluctant to build one, especially if you are a small team with a limited budget.

But there comes a point where fighting lots of local battles with Data Quality becomes more inefficient than building out a framework to reduce Data Quality issues over the long term.

We’re not going to do a deep dive on Data Quality frameworks here, as they are often tied to wider Data Governance frameworks (you can download our guide here) We will say that whatever framework you use, make sure it’s cyclical so that it’s always improving and you are acting on any emerging issues in a timely manner.

How Do I Test for Data Quality?

Classically, Data Quality tests are a set of rules that test between the actual and desired state of data. The desired state may not be perfect, but ‘good enough’. What counts as ‘good enough‘ varies from dataset to dataset, which makes Data Quality more challenging.

What do we normally test in a dataset, though? The DAMA International’s Guide to the Data Management Body of Knowledge says there are six dimensions to Data Quality:

  • Accuracy: does the date look how we expect it to?
  • Completeness: are there any unexpected missing values?
  • Uniqueness: no duplicates!
  • Consistency: does a person’s data match in two different datasets?
  • Timeliness: is the data out of date?
  • Validity: does the data conform to an expected format? Think postcodes, emails, etc.

Tracking all these dimensions for every dataset at every stage of your pipeline is a lot of work, probably too much work. Therefore, a trade-off is often required to focus on areas where Data Quality will have the most impact on the business.

You also have to beware of false positives or minor issues being blown out of proportion, overwhelming your engineers with too many issues. It can help if you categorise your Data Quality issues by severity just like other software issues.

You also have to take into account the mental wellbeing aspect too: few Data Engineers and Analysts want to spend a large percentage of their time fixing Data Quality issues over a long period of time.

There is help, though, with software frameworks to help you write Data Quality testing:

Most of the above profile your data and setup recommended tests for you to use, saving you some time configuring them yourself.

But in reality, we see a lot of custom-made Data Quality testing, partly because Data Quality struggles for investment, so it is usually done in an organic, ad-hoc manner.

The above products work best in development and staging environments, so you can find issues before they enter production or use them as circuit breakers, to stop a Data Pipeline if the incoming or outgoing data is of poor quality.

It is also worth mentioning that you can use constraints in a Warehouse or Lakehouse schema, which have the benefits of not requiring another software library but are not as feature rich (you will likely have to setup your own notifications for alerting).

It is also important to inform your users of any Data Quality issues as soon as possible so they don’t waste time finding out for themselves or use data that is untrustworthy. This can be done through notifications and alerts, though we’ve also had a lot of success creating Data Quality dashboards that sit alongside existing reports and can be easily referred to by users.

Latest Concepts in Data Quality

There has been significant innovation in Data Quality in the last few years, so we present below the concepts to take your Data Quality process to the next level.

This will require more investment, but it will give you an edge over your competitors to make better informed decisions as you’ll have more trustworthy data. This investment should also pay back long term with less time wasted fixing Data Quality issues.

What is Data Reliability and do I Need it?

Data Reliability gives Data Quality more of a support focus, which makes sense as most Data Quality issues in production will be dealt with as a support issue to a Data Platform.

Data Reliability takes a lot of its thinking from Site Reliability Engineering (SRE), which treats support as more of a engineering problem, where you examine your past and current support tickets and look to decrease them with engineering or better processes.

With Data Reliability, you would look to get a baseline of Data Quality issues per week or month and then look at ways to reduce them and monitor to see if the changes have reduced the number of issues and/or reduced the amount of time spent on issues.

The changes to improve Data Reliability can be technology-based:

  • New or updated tooling
  • Better automation of when a pipeline fails or automated actions to respond to a data issue

Or the changes can be process-oriented:

  • Writing better documentation to avoid common issues
  • Incident playbooks so the whole team can more quickly respond to a issue in an consistent way.

You can rather cynically say Data Reliability is just Data Quality with a feedback loop and a time series graph, but it is there to make sure you avoid short term thinking about Data Quality and instead consider long term improvements that will make your data platform more efficient and trustworthy.

Data Reliability Cycle

You may also set targets such as “99.9% of data will refresh on time” or “A maximum of 33% of engineer time should be spent on support issues“ as well. As mentioned before, it can be impossible to achieve perfect Data Quality, so aiming for a reasonable target instead can avoid engineer burnout.

What is Data Observability and do I Need it?

Data Observability is about gaining a Data Platform or organisation-wide understanding of your Data Quality.

It arguably goes beyond Data Quality by adding metadata features normally found in a Data Catalog: cataloguing schemas of datasets and data lineage. These features allow you to more quickly find a Data Quality issue by tracing the lineage of the issue and also you gain the ability to see how much Data Quality is impacting your organisation.

Data Observability software can often also come with Machine Learning (ML) algorithms to detect anomalies in data, so you can be warned about issues you haven’t even thought of yet.

We’ve seen products either extend a Data Quality framework with Data Catalog features such as Monte Carlo and Big Eye. Or existing Data Catalogs add Data Quality functionality, such as Datahub, which imports Data Quality tests created by Great Expectations and dbt tests. Both Soda and Monte Carlo have integration with the Data Catalog Alation.

What are Data Contracts and do I Need Them?

Data Contracts make a contract between a data producer and a data consumer, so the consumer knows what data to expect from the producer.

While you can replicate some of a Data Contract’s benefits by tracking the schema of the data produced, a Data Contract is meant to go beyond that by giving you a full suite of metadata about the data:

  • The data’s schema.
  • How the data is calculated.
  • Who owns the data?
  • What is the data lineage?
  • How to access the data.
  • What is the data’s expected quality, availability, etc.
  • Plus anything else that is relevant to the data.

You may think Data Contracts are redundant if you have a well-maintained Data Catalog, as they capture similar information, but Data Contracts are designed to be checked during every run of a Data Pipeline and have some action in the pipeline if the Data Contract is broken:

  • Stop the pipeline with a circuit breaker.
  • Alerting.
  • Moving data that doesn’t meet the contract to a manual checking table.

For an example, Paypal has open-sourced their Data Contract template.

Data contract schema

https://github.com/paypal/data-contract-template

This should create more positive collaboration between data producers and consumers because they have a collective agreement of what the data should look like. It is not uncommon to have a poor working relationship where a producer makes changes without telling consumers or consumers accessing data in way not recommended by the producer.

One issue with Data Contracts is that they are a new concept, so require more work to implement at present, though that will likely change in the near future as more companies adopt them.

Most of the examples of Data Contracts we’ve seen so far use Apache Flink and the Kafka Schema Registry, so assume you are using streaming, though there are some examples that use batch processing.

Data Governance and Data Quality

Good Data Governance can also improve quality of data. It is important to know where data is coming from, who owns it, for what purpose data is being transformed, and finally, what is the impact of poor availability and data quality: all helped by having Data Governance properly implemented.

Some of the above concepts (Data Contracts and Data Observability) can also improve Data Governance, so investing in Data Quality can also be an investment in good Governance too.

How Does This All Fit Together?

The diagram below is one example of how it all fits together:

  • Any code changes are tested in development and/or test environments with Data Quality Tests to check that any changes won’t have a negative impact on Data Quality.
  • Source Data at the start of the data pipeline is checked to see if the Data Contract is held; if not, a circuit breaker may kick in, stopping the data pipeline early to avoid processing unsuitable data.
  • Data Quality tests are also run in production, which can feel like duplication from testing in development, but there may be changes caused by moving to a production environment (different data, etc.).
  • Data is collected for observability checks by Data Observability software, looking for any anomalous data: a department budget that goes from £10k to £1mil or 10x increase in rows for a table, for example. This can replace a lot of tests, but not all of them.

You’ll also be collecting Data Quality metadata to improve your Data Reliability.

Making all this work together seamlessly isn’t cheap and will take time, but as mentioned, poor Data Quality will also cost an organisation a lot of money. So we recommend tackling this in an agile manner by improving Data Quality in small increments, one change at a time, starting where it will have the most impact.

Summary

Data Quality is a difficult subject to tackle, due to it being a slightly different problem in every organisation and never “perfect”. That said, there are lots of options to help improve the quality of your data, so you should be able to get to “good enough“ if you give Data Quality enough priority and forethought.

Jake Watson is a Principal Engineer at Oakland

Get In Touch 


The post Why Invest in Data Quality? appeared first on Oakland.

]]>
https://weareoakland.com/blog/why-invest-in-data-quality/feed/ 0
How to create a secure Azure Data Platform https://weareoakland.com/blog/how-to-create-a-secure-azure-data-platform/ https://weareoakland.com/blog/how-to-create-a-secure-azure-data-platform/#respond Tue, 01 Aug 2023 09:34:23 +0000 https://www.theoaklandgroup.co.uk/?p=7524 There are many methods you can use to secure your data platform and the data contained within it within Azure. The security controls that will be most effective for each data platform differ based on the usage of the platform, the data sources for the platform and many other factors; Having a holistic view of...

The post How to create a secure Azure Data Platform appeared first on Oakland.

]]>
There are many methods you can use to secure your data platform and the data contained within it within Azure. The security controls that will be most effective for each data platform differ based on the usage of the platform, the data sources for the platform and many other factors; Having a holistic view of the potential options for securing your platform through each of the security layers below will ensure a platform is both fit for purpose and secure.

Diagram

Azure data platform structure

Implementing and maintaining good security within a data platform can be a timely and expensive endeavour so is often overlooked. It can require specialist knowledge to design and implement and make it more complex to connect systems and resources. However, this has to be balanced against the impact of a security breach which could be financial, reputational and have safety implications. During the design phase of a data platform, the sensitivity of data which it will contain, and potential impacts of a breach should be analysed in order to determine an appropriate level of security controls which should be designed into the platform to appropriately mitigate this risk.

Defence in Depth Security:

  1. Data Governance and Classification
  2. Data Protection
  3. Access Control
  4. Authentication
  5. Network Security
  6. Threat Identification and Remediation
  7. Disaster Recovery

 

  1. Data Governance and Classification

Data Governance and Data classification in the context of security is assigning a security rating to data based on the sensitivity of the information contained. Examples of commonly used classifications are ‘Public’, ‘Internal’ and ‘Confidential’. Data in Azure SQL Databases, Azure SQL Managed Instance and Azure Synapse can be allocated a ‘Classification Label’ and an ‘Information Type’. This can be done manually using T-SQL statements or done within the portal within the ‘Data Discovery & Classification’ tab, which can also automatically infer classifications and information types.

Providing classifications to data is beneficial to security as it enables the ability to monitor access to data of different classifications. Additional data governance tools such as Azure Purview can provide additional data classification features such as the ability to associate classifications to data from sources other than those listed above.

  1. Data Protection

It cannot be assumed that data is protected by default even in Azure Platform as a Service (PaaS) services. The nuances in the differences between data protection between services must be understood in order to create a fully protected data platform. Data encryption is an important part of data protection as it protects data from being useable if it is accessed through malicious activity. Many Azure services provide a certain level of data encryption by default, but this cannot be assumed to be true. Azure SQL Database and Azure SQL Managed Instance both have Transparent Data Encryption (TDE) enabled by default, this service is also available for Azure Synapse Analytics Dedicated SQL Pools, however in this case it is not enabled by default. Given the rise in popularity of lakehouse based architectures within data platforms, it is also worth noting that Azure Data Lake Storage (ADLS) also utilised encryption at rest and in transit, to secure underlying data.

  1. Access Control

Access to data and resources should be granted using the principal of least privilege. So users are only granted access they need to perform their duties and no more. Access should be regularly reviewed and when no longer required it should be revoked. Azure tenants should be carefully designed with Management Groups, Subscriptions and Resource Groups to reduce unnecessary resource visibility, for example preventing the Marketing Department from seeing or accessing Finance Department resources.

Typically the fewer people who have access to a resource or data the more secure it is, this limits the chance of users accidently or deliberately leaking or altering potentially sensitive or business critical data. It’s not only access to data which should be carefully considered, access to resource configuration is also important. For example, if a user is granted contributor access to a resource they could delete or alter the resource by mistake. Additionally, they could make changes which compromise the security of the platform, such as altering or removing Network Security Group (NSG) rules without understanding the consequences, permissions like these should be limited to those with requirement for it and required technical understanding.

Sensitive data can be protected from unauthorised access through the use of data masking. This feature is available to Azure SQL Databases, Azure SQL Managed Instance and Azure Synapse Analytics. The amount of data which can be viewed but different users or user groups can be dynamically defined using policies. For example a Database can be configured so that members of the Finance team are able to view full credit card numbers, HR personnel are able to view the last 4 digits of the card number and developers are only able to a series of ‘X’s. The dynamic data masking feature on the supported services listed above can be configured using T-SQL statements or within the Azure portal on the ‘Dynamic Data Making’ page.

Resource locks can be applied within Azure to control which users can perform certain operations on resources. There are two types of resource locks: ‘Read-Only’ and ‘Delete’. ‘Read-Only’ locks allow users with access to view a resource but they cannot make any changes to it. ‘Delete’ locks stop users from being able to delete resources. These features prevent accidental resource deletion or reconfiguration.

  1. Authentication

Authentication is the method by which users or services verify their identity when attempting to access another service. Azure services have support for authentication and access control with Azure Active Directory (AAD). AAD offers many features to enhance security such as Single Sign On (SSO) which allows users to use one set of credentials to sign on to multiple services.

By default, Role Based Access Controls (RBAC) permissions on resources are granted with authentication using AAD, which can be done for individual users or users grouped into management groups. Data access can also be authenticated using AAD with support for Multi Factor Authentication (MFA) through the use of Azure SQL Database, Azure SQL Managed Instance and Azure Synapse Analytics. This removes the need for additional passwords to access SQL and reduces the likelihood of many users signing in using the admin credentials when this is not required. Also integrating SQL access with AAD means if employees are removed from the organisations AAD for example due to leaving the company then their access to SQL will also be automatically removed without needing to manually delete their SQL user or rotate the password of a shared login.

Authentication is required between services and resources, traditionally this authentication is done using a username and password combination, however this is vulnerable to these credentials being leaked. A more secure method of authenticating between systems is using Managed Identities. When using Managed Identities in Azure, the identity of a resource or service is registered within AAD and other services can use AAD tokens to authenticate, rather than  a username and password.

  1. Network Security

An appropriately designed network security framework in Azure can protect a data platform from attack and unauthorised access, whilst also allowing for functional communication in and out of the network. Features offered by Azure to create a secure and functional network topology include private endpoints to secure PaaS services, Azure firewall and virtual networks with associated NSGs.

Virtual networks create groups of connected services which can be protected from unwanted inbound and outbound communication. Virtual networks are split into subnets, these subnets can have associated NSGs which are lists of rules which allow or deny traffic. When assigning NSGs all communication should be blocked as a default and only exceptions for specific purposes should be made in order to minimise communication. Reducing the number of allowed protocols and allowed ports for communication on your virtual network reduces the routes an attacker could take into your network.

Private endpoints within Azure are available for a range of PaaS services such as Azure Data Lake Gen 2 and Azure SQL Database. A private endpoint creates a Network Interface Card (NIC) which is associated to the virtual network. Then public access to these services can be disabled and only traffic to and from the network can be allowed, creating a private connection. The use of private endpoints can reduce the risk of data leaks as data is always contained within the private network and is not transferred publicly over the internet.

Azure Firewall can be used to inspect and analyse traffic coming from outside an Azure environment into it and between spokes within an Azure hub and spoke model. The firewall can inspect and block unwanted traffic to your Azure environment. Azure Firewall is integrated with Azure Monitor, Azures logging and alerting offering, so metrics from the firewall can be inspected.

Many data platforms require connectivity to on premise systems, these connections can pose a potential security risk of not properly configured and secured. There are several methods for connecting cloud and on premise systems within Azure. When choosing what method to use to connect systems the cost of the connection must be balanced with the required security for the connection. Azure ExpressRoute can provide a dedicated private connection between an Azure environment and an on premise system so no data transferred is exposed to the internet hence making this a very secure option. However Express Route is a costly option. Another option for creating this connection is to use a Site to Site Virtual Private Network (VPN) which transfers encrypted data between the on premise system and Azure over the internet. A Site to Site VPN is typically cheaper than Azure Express Route, but due to data being transferred over the internet it is considered to be less secure.

  1. Threat Identification and Remediation

It can often be difficult to identify security threats to a data platform, you can collect logs and query them in order to identify suspicious or threatening activity but it can often be difficult to filter out unnecessary logs and identify useful information. The Microsoft feature ‘Microsoft Defender for Cloud’ provides detailed recommendations of how resources can be configured to be more secure and of active cyber security alerts such as logins from unusual locations. Defender for Cloud also provides recommendations for remediation steps which should be completed to resolve potential security issues.

Most resources within Azure have the ability to export diagnostic logs and metrics to Azure Monitor or Azure Log Analytics Workspace which are central repositories where logs can be queried and alerts can be set based upon these logs. For example security logs can be exported from Virtual Machines which detail login attempts, then Azure Log Analytics Workspace can be used to query failed login attempts and an alert can be created to email nominated users when a failed login attempt has been made.

  1. Disaster Recovery

In a worst case scenario, a cyber attack (or accidental deletion) could result in services and data being deleted or un recoverable. In this scenario a robust and timely disaster recovery plan which minimises the Recovery Point Objective (RPO) is imperative. This can be achieved through a variety of methods. Storing templates for infrastructure deployment, and ARM templates for orchestration services such as Azure Data Factory in Repos will mean a platform can be redeployed quickly and easily.

Many Azure PaaS storage services have features to simplify recovering lost data which are enabled by default. For example Azure Synapse Analytics creates regular restore points on dedicated SQL pools and a deleted or corrupted database can be redeployed from these restore points.

Azure provides geo-replication features to protect resources and data against a potential disaster at one of its data centres or regions. If enabled on resources geo-replication can create a replica of the resource within a different availability zone or region so if the hardware containing the primary resource is damaged the resource is still available through the replica. This feature is available on many resources in Azure including Azure SQL Database and Azure Storage Accounts.

Abigail is a Senior Data Engineer at Oakland

The post How to create a secure Azure Data Platform appeared first on Oakland.

]]>
https://weareoakland.com/blog/how-to-create-a-secure-azure-data-platform/feed/ 0
What are the challenges of building a data platform? https://weareoakland.com/blog/what-are-the-challenges-of-building-a-data-platform/ https://weareoakland.com/blog/what-are-the-challenges-of-building-a-data-platform/#respond Tue, 14 Mar 2023 16:27:32 +0000 https://www.theoaklandgroup.co.uk/?p=7122 When building a data platform for your business, you need to anticipate and plan for any potential challenges you may encounter. That’s where a data engineering specialist can help.  With decades of data experience behind us, we’ve seen and dealt with practically every problem you might encounter during data platform development. In this guide, we’ll...

The post What are the challenges of building a data platform? appeared first on Oakland.

]]>
When building a data platform for your business, you need to anticipate and plan for any potential challenges you may encounter. That’s where a data engineering specialist can help. 

With decades of data experience behind us, we’ve seen and dealt with practically every problem you might encounter during data platform development. In this guide, we’ll cover common difficulties you might face when developing your data platform, with advice on how to handle each to make your data platform journey as smooth as possible. 

To understand more about the complexities of a data platform and the process of building one, check out our data platform guide.

What is a Data Platform?

Put simply, a data platform is a central hub that collects, organises, transforms and applies your data. You can learn more about data platforms through our expert data platform guide

What Challenges Can Impact Building a Data Platform?

It would be lovely if everything could be easy, wouldn’t it? But unfortunately, many of the best things in life come with a few struggles along the way to attain them. Luckily, our data engineering experts can help you navigate the following common challenges.

Challenge 1: Excessively Tech-Centric Focus 

It’s easy to think of your data platform initiative as a technical project; after all, you’ll soon be designing, launching and modifying an expensive chunk of technology real estate. 

But think back to the failed data management initiatives you’ve observed in past organisations – what did they have in common? Chances are, they got bogged down in the tech aspect of building a data platform at the expense of the business strategy. 

Technical teams often prefer to solve technical problems rather than get involved in the messy business of persuading people with different objectives to collaborate. 

The big risk in being overly tech-focused is that if your data platform does not meet user needs (such as data availability and usability), users won’t use it. You will achieve minimal adoption, and the platform will fail to deliver tangible business outcomes. 

Therefore, your data platform aspirations should form the ‘pointy end’ of a data strategy – it’s where the rubber hits your digital transformation roadmap. 

Whenever you feel the narrative swinging too far over to the tech, bring it back with questions such as: 

  • What does this tech mean to our business model and value proposition? 
  • How will the technical direction impact our strategic goals? 
  • How will the tech impact the customer experience? 

Without this, you run the risk of low adoption, low involvement from business users, and ultimately, low value delivered, if any. 

Challenge 2: Departmental Data Silos 

In an attempt to solve a tactical or near-term challenge, departments or cross-business functions can often be swayed by a solution vendor’s shiny offering. The department then commissions a localised solution that seemingly fits their needs but doesn’t take stock of the wider data strategy or business needs. 

Other teams then struggle to extract this new data, particularly if the department has used off-the-shelf solutions. 

The result is an ever-increasing technical burden that becomes difficult to unravel and migrate in the future. 

To prevent this from happening, it’s crucial to implement solutions in a consultative and collaborative manner. What do you need from your solution? Is there one out there that can benefit a greater range of stakeholders? 

Challenge 3: Poor-Quality Data

One of the most common issues we have to deal with is poor-quality data. This is important since your data platform will only deliver your business goals if the data ingested is high enough in quality otherwise your platform could sink without a trace.

But what is data quality? Data quality can be measured by the following six principles, which you should always keep in the back of your mind when building a data platform: 

  • Accuracy: Does your data correctly represent the intended entities and events? Are your data sources reliable? 
  • Consistency: Is your data uniform across systems and data sets?
  • Validity: Does your data conform to predefined rules regarding data structure and values?
  • Completeness: Have you included all the expected data values and types?
  • Timeliness: Is your data current and available to use when needed?
  • Uniqueness: Have you ensured no duplicate records within a single data set? Can your data record be uniquely identified?

To aid you in understanding how your data measures up to these principles, Oakland offers data quality assessments with advice on improvement as part of our data governance service

Why is data quality important? Poor quality data leads to inaccurate analytics, which, in turn, leads to bad business strategies. This can result in missed sales opportunities, decreased customer satisfaction, significant monetary fines over incorrect reporting and a lack of trust from your corporate executives and business managers. See our blog on why you should invest in data quality to see the benefits it can provide for your business. 

Challenge 4: Lack of Data Governance 

According to Gartner, “Data governance is the specification of decision rights and an accountability framework to ensure the appropriate behavior in the valuation, creation, consumption and control of data and analytics.”

The demand for data governance originally emerged from the shift toward more robust regulatory controls in the banking and insurance sectors. 

Today, data governance is pervasive across all industries, and now you can even buy data governance platforms off the shelf! Yet, many data platform initiatives stutter or fail when Data Governance is immature or lacks key components suited to Data Platform strategy and management. 

The impact of poor data governance can include: 

  • Business and technical users lose trust in the data 
  • Designing and maintaining the platform takes far longer 
  • Frustrating ‘Turf wars’ over data ownership erupt or remain unresolved 

To find out more about data governance and how you can enable it properly when designing a data platform, check out our guide: How to Launch a Data Governance Initiative by Stealth

Challenge 5: Failing to consider the complexity and cost implications of the legacy data landscape 

Your data platform is not an island; it needs careful integration with existing systems and processes. 

At Oakland, we’ve been around a long time and one of the recurring trends we’ve seen is the case of the ‘over-optimistic’ target vendor. 

Despite the glossy marketing blurb, no data platform is a true plug-and-play solution, so avoid anything that sounds too simple to be true. The last thing you want is to get locked into a vendor, meaning that you are forced to keep using a low-quality platform because switching would be impractical. 

Many aspects of vendor lock-in are inevitable, and using a vendor’s portfolio of cloud-native services increases lock-in. However, access to integrated services and increased discounts can be a plus. To prevent lock-in (e.g. open source software solutions), make decisions on a case-by-case basis and vet the total cost of ownership (TCO) of solutions intended. 

Creating a new data platform in any enterprise requires a careful analysis of what approaches have gone before and now require direct integration with your new data architecture.

We’ve parachuted into several data platform recoveries in which the complexity of integration was overlooked and soon became the mother of all obstacles to going live. 

In short, don’t overlook the essential brownfield discovery tasks that some vendors like to gloss over to get you over the finishing line. 

Challenge 6: Not prioritising requirements 

As you get deeper into data platform delivery, stakeholders asking, “So, what are we building again?” can become a common challenge. 

The problem is the modern data platform strategy can support a range of use cases, including: 

  • BI (Business Intelligence) /MI (Management Information) capabilities 
  • Self-service reporting 
  • Integration 
  • Data catalogue
  • Customer Data Platform, e.g. single voice of the customer, employee, product, service or asset 
  • Executive or regulatory reporting 
  • Advanced analytics and AI 
  • Real-time analytics 
  • Assessment management 
  • Adjunct to a core system.

It’s easy for your data platform to lack clarity and prioritisation around its core function, especially as different groups begin to see it as a data ‘dumping ground’. The practice of “let’s keep it in case we need it” can lead to a bloated data platform, further complicating the task of extracting insights from your data. 

To prevent a toxic swamp of data, we prefer to phase the delivery of a data platform with a regular cycle of ‘Lighthouse Projects’ that solve burning issues within the business but still align to an overarching data strategy and architecture with a clear transition to an enterprise solution. 

Start small, think big, and act fast. 

You get to demonstrate the benefits of each release, garnering support as you deliver each successful project, and helping justify further initiatives’ investment. 

Challenge 7: Creating the case for change 

Creating a compelling business case for a modern data platform can be challenging for many organisations, particularly when faced with a legacy of delivery struggles. 

Deploying the next generation of data platforms has multiple benefits and use cases that will appeal to various leadership sponsors. 

We’ve found that the key is smaller, faster pilot projects that deliver rapid and sustained gains without over-investment and risk while building data capabilities at the same time. You can find the right use cases for your pilot project with Oakland’s complementary use case workshop, where we’ll explore what the business wants to know and where insights are needed. 

In short, focusing on delivering the right data at the right time to support specific business outcomes will quickly gain support for the future of your data platform.
Build your perfect data platform with Oakland and explore our other services to take your data even further. You can also contact us with all your burning data platform questions. Want to learn more about what we do? Take a peek at our blog or give our podcast a listen.

The post What are the challenges of building a data platform? appeared first on Oakland.

]]>
https://weareoakland.com/blog/what-are-the-challenges-of-building-a-data-platform/feed/ 0
What is a data platform? https://weareoakland.com/blog/what-is-a-data-platform/ https://weareoakland.com/blog/what-is-a-data-platform/#respond Thu, 02 Feb 2023 12:34:30 +0000 https://www.theoaklandgroup.co.uk/?p=6904 In today’s digitised business landscape, the ability to collect, analyse, and manage data effectively can be the critical difference between a business’s success and failure. But what exactly enables businesses to harness the full potential of their data? Enter the customer data platform. In this blog, we’ll discuss everything you need to know about data...

The post What is a data platform? appeared first on Oakland.

]]>
In today’s digitised business landscape, the ability to collect, analyse, and manage data effectively can be the critical difference between a business’s success and failure. But what exactly enables businesses to harness the full potential of their data? Enter the customer data platform.

In this blog, we’ll discuss everything you need to know about data platforms, including their many benefits and how they can improve your business. 

What is a Customer Data Platform? 

A customer data platform is a centralised system for collecting, storing, and managing data from various sources. It uses innovative data engineering to leverage components such as data storage, processing, analysis, and integration. This lets organisations gather insights, make data-driven decisions, and ensure data quality and security. 

What is Data Engineering?

Data engineering focuses on the practical aspects of data collection and data analysis. It comprises the design, development, and management of a central system that facilitates the collection, storage, and processing of large volumes of data. 

Key aspects of data engineering include data collection, processing, performance optimisation, security, and seamless collaboration with other teams.

What are the Benefits of Data Engineering?

Data engineering offers a number of handy benefits that can significantly improve a business’s efficiency. The main notable benefits include:

  • Improved Data Quality: Data engineers ensure that the data used for analysis is accurate and reliable through rigorous cleaning and transformation processes.
  • Enhanced Efficiency: Automated data pipelines reduce the time and effort required to manage data, allowing for faster access to insights.
  • Scalability: Data engineering solutions are designed to handle growing volumes of data, ensuring that systems can scale with the business.
  • Better Decision-Making: Data engineering supports data-driven decision-making across the organisation by providing clean, integrated, and accessible data.
  • Cost Savings: Efficient data management and storage solutions can reduce costs associated with data handling and infrastructure.

To learn more about further advantages, visit our dedicated guide: Why do you need a data platform? 

What are Cloud-Native Platforms?

Cloud-native platforms are essential tools to help accelerate the execution of plans over the next 2-3 years. Improved access to cloud services enables the introduction of modern technologies with less operational burden than legacy systems, expediting and facilitating the creation of innovative business solutions.

Adopting cloud-native platforms provides the primary means for enterprises to execute their digital strategies, enabling business growth, customer retention, and efficiency. According to Gartner, cloud-native platforms will serve as the foundation for more than 95% of new digital initiatives by 2025, up from less than 30% in 2021.

What are the Key Components of a Data Platform?

There are several elements that collectively make up a successful digital data platform, including:

Data Access and Governance

In simple terms, a data platform enables data access, governance, delivery, and security. It brings together the technology needed to collect, transform, unify, and govern the data required to support users, applications, models, and data products.

Scalability and Security

To survive in today’s fast-moving market, an organisation’s data platform must be cost-effective, highly scalable, and secure from the outset. 

It should enable data ingestion from multiple sources, including other data platforms, and be flexible enough to accommodate system changes in the future. What’s more, the architecture of your data platform must effectively support your business outcomes.

The Role of Data Governance

A successful data platform combines cloud capabilities with a robust data governance approach. Without data governance, issues like poor data quality and availability can persist, impacting current performance and hindering future growth opportunities.

Targeting Organisational Needs

A data platform must address an organisation’s specific needs. For example, a complex organisation with many data sources will need data conformity as a critical capability. 

We see a data platform as a layered set of capabilities that build on each other. This enables organisations to realise value from high-quality data managed through governance processes, thereby empowering confident decision-making.

The Importance of Cloud-Native Solutions

We believe any modern-day data platform should be cloud-native.

Cloud-native platforms allow organisations to deliver scalable solutions without heavy reliance on managing the underlying infrastructure. These platforms are typically sourced from public cloud services (e.g., Amazon Web Services, Microsoft Azure, Google Cloud Platform) or created using software that structures a private cloud environment for added security and control.

What are the Core Capabilities of Cloud-Native Platforms?

Cloud-native platforms use core functionalities such as container management, infrastructure-as-code, and serverless functions while supporting continuous integration and delivery pipelines. These platforms can work with other cloud tools, SaaS tools, or on-premise applications, offering a speedier alternative to traditional on-premise solutions.

Enabling core capabilities (shown in the diagram below under the themes of knowledge, insight, and awareness) can help realise the true value of the cloud through the provision of all three layers— infrastructure, services, and governance.

As you seek to progress from knowledge to insight and awareness, your capabilities must evolve to meet your ambitions. The component parts of a data platform are shown below.

Data Journey 

How to Build a Data Platform

The easiest way to visualise building a data platform is to think of it in terms of building a house. 

When building a house (data platform), you don’t just start laying bricks (processing data). You need to know the room measurements (data subject areas), layout (data models), and adherence to building regulations (governance). This is how you design a data platform that delivers on your business goals.

For more expert knowledge on building a data platform, visit our blog: What are the challenges of building a data platform?

What are the Benefits of Cloud-Native Platforms?

As technology advances, we’ve seen a number of evolving benefits to opting for a Cloud-Native platform for your data strategy. These include: 

  • Reduced Total Cost of Ownership: Cloud data platform deployment and operation costs are generally lower than traditional on-premise implementations.
  • Service Resiliency and Management: Cloud data platforms require less effort and complexity to meet spikes in demand while maintaining operations.
  • Speed and Service Agility: Cloud-based platforms significantly reduce delivery timescales, leveraging reusable cloud blueprint architectures and components.
  • Business Model Transformation/Optimisation: Cloud-native platforms support broader benefits beyond cost and speed, enabling business model transformation and optimisation.

With O’Reilly survey data showing that 77% of organisations were already cloud-native or pursuing a cloud-first strategy in 2021, the benefits are already being understood across industries.

Discover More Data Strategy Solutions with Oakland

Ready to transform your business with a modern cloud-based platform? Download our guide to delivering one, or contact us if you have any questions.

Want to stay up to date with the latest industry insights on all things data? Explore the Oakland blog today.

The post What is a data platform? appeared first on Oakland.

]]>
https://weareoakland.com/blog/what-is-a-data-platform/feed/ 0
Does Microsoft Purview solve the Data Governance Challenge? https://weareoakland.com/blog/does-microsoft-purview-solve-the-data-governance-challenge/ https://weareoakland.com/blog/does-microsoft-purview-solve-the-data-governance-challenge/#respond Tue, 22 Nov 2022 13:11:36 +0000 https://www.theoaklandgroup.co.uk/?p=6855 Microsoft Purview Data Governance for the cloud, on-premise, multi-cloud and office 365 workloads. Introduction Most organisations are exploding with data that has been collected, transformed, and reported on, but this data is often not well-tracked as the organisation becomes more data-driven, increasing two pain problems that have been growing for the last few decades: How...

The post Does Microsoft Purview solve the Data Governance Challenge? appeared first on Oakland.

]]>
Microsoft Purview

Data Governance for the cloud, on-premise, multi-cloud and office 365 workloads.

Introduction

Most organisations are exploding with data that has been collected, transformed, and reported on, but this data is often not well-tracked as the organisation becomes more data-driven, increasing two pain problems that have been growing for the last few decades:

  • How can we audit all this data to protect against data leaks and unexpected data loss?
  • How can data users discover data in an environment that changes constantly?

Data Governance Products help mitigate these problems, among others, but are often complex due to requiring:

  • The ability to scan a large variety of data sources
  • A highly customised user interface
  • A powerful search engine to find data assets by many different types of metadata attributes

These are just some of the main requirements that create a software marketplace full of products that are often expensive and hard to implement and maintain.

These products also need to ingest large amounts of sensitive organisational data to meet user requirements, ironically creating a Data Governance concern in itself!

Purview aims to ease the pain of Data Governance by being feature-rich, easy to deploy, maintain and secure. But is it worth the cost, and can it compete with bespoke Data Governance companies that have a head start measured in years or even decades?

Purviews Features

  • Its connectors are very Microsoft-focused but cover most of its ecosystem: Azure, SQL Server, Power BI, and Office 365. If you’ve already bought heavily into Microsoft, you can scan most or all your data assets automatically.
  • It focuses less on connectors made by other companies but still covers many popular data products like SAP, Salesforce, Oracle, GCP Big Query, AWS S3, and Snowflake.
  • It offers a lot of flexibility in managing Data Catalog users with 9 different roles to choose from and syncs up to your Azure Active Directory groups and users.
  • Can classify data with 200+ pre-built classifications, as well as custom classifications.
  • Business Glossary with an extensive text editor. Ability to add contacts for roles like Data Owner and Steward to each data asset.
  • Pre-built reports to quickly check insights such as what percentage of data has a Data Owner and the percentage of new data assets in the last month.
  • Offers an API and Python SDK for making custom data sources where connectors don’t exist or mass updating existing scanned data assets.
  • It doesn’t offer much insight into Data Quality of scanned data assets, which can be found in other Data Governance products. However, it could theoretically push Data Quality metrics to Purview via its API.
  • Data Sharing allows users to give other users read-only Data Lake data access without having to copy data.
  • Can ingrate Data Governance with Master Data Management using Profisee
  • Purview is relatively new after only being available to customers for a few years but it is receiving heavy investment from Microsoft, with new features appearing monthly.

Deployment

  • As someone who has designed and built many data platforms, I highly value any product that can be deployed quickly, has low maintenance, and will meet strong client IT & security requirements. I believe Purview is stronger than most Data Governance products in this area.
  • It is as easy to deploy and maintain in Azure as any SaaS data governance product but also offers a choice – 20 plus regions to deploy into, including the UK.
  • Purview can also keep all traffic in and out of its server on its private network using Private Endpoints, never touching the public internet, offering an extra layer of data security when creating a Data Catalogue.
  • Can scan Azure Data Products via Managed Identity authentication offering high-security data connections without worrying about managing passwords.
  • It can connect directly to scan on-premise and other cloud data assets, though it requires some technical knowledge.

Costs

Automated Data Governance tooling is often not cheap, with costs starting in the thousands of pounds for most products. Purview is no exception: it has a base price of £250 per month and will cost more if scanning large workloads. However extra capacity is costed using the pay as you go model, so you only get charged extra when scanning lots of data.

There are also additional extra costs depending on which features are used.

Due to the pricing being highly variable in Purview, many organisations will build a Proof of Concept to road-test Purview for a month or so to accurately measure costs.

Alternatives

Note this isn’t a comprehensive list and is a quickly evolving space with new exciting start-ups entering all the time.

  • Build your own:
    • Excel: low cost, low maintenance if data structures don’t update regularly, doesn’t require any specialist skills to build. While we suspect this is the most common type of data catalogue used, we feel nervous about doing a Data Catalog in a data tool infamous for having poor Data Governance. It does not scale and requires lots of manual effort to work out data lineages and classify data sensitivity.
    • Automate your own solution by extracting schemas of databases and files. This is a nice quick way of generating a data catalogue with low maintenance and little extra costs. You can also build a dashboard on top of the Business Intelligence (BI) platform of your choice. It requires minimum effort if the number of data assets is small, though adding features like data lineage and classifying data sensitivity will require a reasonable amount of engineering effort.
  • Databricks Unity Catalog – ideal for Databricks heavy data platforms, as it is a free extra. Though it will only scan what Databricks can scan. You can integrate with other Data Governance products and update schema as they update in real time, which you don’t see much in other Data Governance products.
  • Other clouds Data Governance solutions like AWS Glue Data Catalog and GCP Dataplex. Both are arguably less feature-rich than Purview, especially for non-technical users, though they are easier to implement if most of your data assets are in their respective clouds. Also, both are used to ingest into other larger Data Governance products.
  • Mature products like Informatica and Talend. These tend to charge by the user and are more commonly found on-premise (though they can be configured and maintained in the cloud on Virtual Machines). They will likely cost the most; sometimes, this is significant, but these are well-trusted and reliable. Add the most value if you buy into the rest of their data platform ecosystems.
  • New products like Atlan and Immuta often focus on providing Data Governance to more recent data products like Databricks and Snowflake but also often focus on making deployments into the cloud more accessible by offering deployments via Docker or Kubernetes. Immuta also provides a single pane of glass for fine-grain data access across many popular data products that allows data access controls at a column and row level.
  • Open Source software like Datahub and Amundsen, both built by large tech companies (LinkedIn and Lfyt respectively). These are the go solutions if your organisation has the technical capacity to build and maintain complex workflows. They offer a lot of customisation and no licence costs, so they can be much cheaper at scale and be more custom tailored to fit an organisation’s Data Governance needs.

Summary

In an increasingly challenging data governance market, Azure Purview is a serious option to consider despite missing some features compared to more bespoke Data Governance companies.

However, using Purview if you spend hundreds per month or more on Data Governance is not a good return on investment, or you want a solution that easily integrates Data Governance and Quality.

If you are looking for a Data Governance product that is easy to deploy, secure, catalogue, and classify data assets, and provides some customisation through APIs and user interface at a competitive cost, then we think Purview is a good contender.

Jake Watson is a Senior Data Engineer at The Oakland Group

The post Does Microsoft Purview solve the Data Governance Challenge? appeared first on Oakland.

]]>
https://weareoakland.com/blog/does-microsoft-purview-solve-the-data-governance-challenge/feed/ 0
Taming your data assets with Databricks https://weareoakland.com/blog/taming-your-data-assets-with-databricks/ https://weareoakland.com/blog/taming-your-data-assets-with-databricks/#respond Mon, 14 Nov 2022 20:25:45 +0000 https://www.theoaklandgroup.co.uk/?p=6829 Safely, securely, and efficiently handling data at any scale is challenging. Here at Oakland, we’ve had years of experience helping complex organisations tame their vast data assets to draw meaningful insights from them. These years of experience and the fact we are passionately tech-agnostic enable us to recommend the right tool for the job. One...

The post Taming your data assets with Databricks appeared first on Oakland.

]]>
Safely, securely, and efficiently handling data at any scale is challenging. Here at Oakland, we’ve had years of experience helping complex organisations tame their vast data assets to draw meaningful insights from them. These years of experience and the fact we are passionately tech-agnostic enable us to recommend the right tool for the job.

One of the most impressive tools in our kit bag is Databricks – which has suited our needs in the data landscape for three main reasons: power, flexibility, and a low barrier to entry. What Databricks is has been covered https://hevodata.com/learn/what-is-databricks/ extensively https://medium.com/codex/what-is-databricks-and-how-can-it-be-used-for-business-intelligence-6ac62cac198a , but a significantly more interesting question is why it has been so widely adopted?

A History Lesson

Before 2000, the dominant form of data storage was the Relational Database Management System, which was generally interrogated and created with SQL. However, post-millennium, the huge rise in web traffic created a vast quantity of semi-structured and unstructured data. Everything could be recorded: clicks, user patterns, tweets, purchasing patterns, sound, and video. But that was just the beginning; with the rise of IoT devices and the proliferation of cheap telemetry, this paradigm of collecting a large quantity of dissimilarly structured data is not abating– and is one of the most significant problems we help our clients with here at Oakland.

We have increasingly seen enterprises turn to the Data Lake to store this data. Unlike an RDMS where data is stored neatly in related tables, a Data Lake is effectively an open storage medium either on-premises or in the cloud. An organisation can pour its data into this Lake to be retrieved, structured, and analysed later. The first and most apparent problem is organisation: the early Data Lakes did not have folder structures. Even now, the supported folders in AWS’s S3 buckets are a naming convention only. A more pressing problem, though, was of scale.

With growing quantities of data, analytical operations become increasingly difficult. It becomes impossible to load all the data into a single computer’s memory simultaneously. Even if that were possible, the time taken to perform even simple aggregations or analysis was becoming unacceptably long. Parallel operations are required, where the data is broken up into several chunks and operated on by a cluster of processors – all of which are orchestrated by a central driving processor.

The Power and The Storage 

To offer anything to the Big Data marketplace, Databricks would have to leverage the power of parallel computing. Databricks sits on top of Spark, which is Apache’s open-source engine for Big Data analysis. More than that, their CTO was the creator of Spark, and Databricks remain significant contributors to the codebase. Databricks also have some closed-source optimisations which can only be accessed through the product itself. With those improvements, Databricks claim processing speeds of several times faster than the bare Spark product.

To properly leverage Spark, Databricks runs on a cluster of compute resource. The amount of memory and number of CPUs the cluster has are configurable – with a balance to be made between the speed of queries and the cost of maintaining the cluster.  Also with this, Databricks offer several different runtimes – which are the sets of core components running on the compute resource. Runtimes are picked depending on the use case, whether general purpose, ML specific or tailored to intensive SQL queries. Importantly though, leveraging this power is easy for the users of the platform. After an initial configuration, data professionals can continue their work.

At Oakland, one of the things we think makes us different is that we pride ourselves in doing work that sets clients up for success in the future. The ease with which Databricks’ power can be stood up and maintained is a key factor in why we use it and recommend it to our clients.

Alongside the processing power to run analytical workloads, Databricks workspaces also include managed storage. In particular, their ‘Delta Lake’ paradigm allows for the retention of much of the flexibility of a Data Lake while retaining many of the advantages of a traditional relational database. The Delta Lake is built on parquet files, which are accompanied by metadata JSON recording all changes (deltas) made to that file. As a result, ACID transactions are possible, and there’s a clear governance record for each piece of data.

Low Barrier To Entry 

Another factor in our choice to implement Databricks can often be its incredibly low barrier to entry for developers to work on the platform. The central component of the workspace which most analysts or developers in any organisation will interact with is the notebook. Heavily inspired by the Jupyter notebook – developing inside it should be familiar to most coding data professionals. There’s also flexibility in the language of development. Python, R, Scala, and SQL are all first-class citizen languages in Databricks – though the Spark bindings of Python and Scala make them arguably the most powerful of the four. This familiarity with both environment and language speeds up both onboarding and development time, adding weight to the choice of Databricks.

We have also found Databricks to have a low barrier to entry from the infrastructure and set-up side. While at Oakland we have created and maintained complex Databricks platforms, it is also possible to create them graphically from a cloud portal – or with a small amount of Infrastructure-As-Code.

Flexibility 

The final reason for our clients to choose Databricks is the platform’s flexibility. At the end of the day, the Databricks notebook is effectively just some arbitrary code running on a Linux machine – so it can, in theory, contain any process imaginable. Orchestrating and automating the running of whole notebooks with Databricks Jobs is a potent tool – and one we have taken advantage of multiple times here at Oakland for our clients. With notebook automation, Databricks can stand up as nearly any part of an enterprise’s ETL process.

We have used Databricks as an ingestion engine to drive the aggregation of many disparate data sources into one Data Lake. This leverages the power of Databricks jobs – running a notebook to pull data from many sources every few minutes. We have also found Databricks helpful on the analytical side when clients have asked us to use notebooks to create custom aggregations and transformations to feed BI dashboards. It has also been possible to use the notebooks as dashboards with widget visualization functions – even as a quick measure to demonstrate to clients some initial exploration of their data.

With a global recession on the horizon, organisations will be looking to drive valuable insights from an ever-increasing pool of data; quickly and democratically. Here at Oakland, we are seeing increased use of tools like Databricks for the reasons we began this blog with; power, flexibility, and a reasonably low barrier to entry. If you’d like to speak to one of our tech team about how Oakland has successfully implemented Databricks, please email hello@theoaklandgroup.co.uk

Mike Le Galloudec is a Data Engineer at The Oakland Group

 

 

 

 

The post Taming your data assets with Databricks appeared first on Oakland.

]]>
https://weareoakland.com/blog/taming-your-data-assets-with-databricks/feed/ 0