Platform & Data Engineering | Oakland Fri, 09 Jan 2026 09:30:35 +0000 en-GB hourly 1 https://wordpress.org/?v=6.9.4 https://weareoakland.com/wp-content/uploads/2024/01/cropped-oakland-favicon-150x150.jpg Platform & Data Engineering | Oakland 32 32 Common Data Challenges & How to Avoid Them https://weareoakland.com/blog/common-data-pitfalls/ Thu, 30 Oct 2025 08:31:19 +0000 https://weareoakland.com/?p=9782 Data Decoded: Calling out some of the traps we see in data programmes. The last blog in my ‘data decoded’ series looks at the common issues we see when data platforms are built and automatic ROI is expected. Familiar Data Platform Challenges & Downfalls Department A builds a platform, Department B builds their own. Two...

The post Common Data Challenges & How to Avoid Them appeared first on Oakland.

]]>
Data Decoded: Calling out some of the traps we see in data programmes.

The last blog in my ‘data decoded’ series looks at the common issues we see when data platforms are built and automatic ROI is expected.

Familiar Data Platform Challenges & Downfalls

  • Data Silos

Department A builds a platform, Department B builds their own. Two platforms that don’t talk = fragmented view, no enterprise coherence.

Explore this further in: Should You Build or Buy Your Data Platform?

  • Technology First Mindset

“Let’s build a lakehouse then figure out what we’ll do with it.” Wrong order.

  • Lack of Business Alignment

If business users aren’t engaged, platforms sit idle.

  • Ignoring Usage / Adoption

You may have dashboards, but no one uses them.

  • No Clear ROI Tracking

If you cannot measure value, you cannot grow investment.

  • Over-Engineering Early

Streaming, ML, fancy stuff before you’ve nailed decision-making and adoption.

Avoid these data challenges by starting small, focusing on decision-making, involving business users, measuring value, and building iteratively.

The Role of the Data Platform Consultant

As a data platform consultant, you’re not just building technology – you’re helping your client (or your business) become decision-driven. Your role is to:

  • Ask ‘why?’ early and often: Why are we building this? What decision will it support?
  • Translate business outcomes into data product features: What dataset, what transform, what serve layer?
  • Build lean-first: Minimum viable platform that supports decisions, then scale.
  • Embed metrics & governance: Usage analytics, data quality, cost control.
  • Communicate value: Demonstrate ROI, show wins, secure funding for next phases.

In short, you’re a bridge between the business (who must decide) and the engineers (who build) – the ultimate solver of data challenges.

A Realistic Journey

Reporting → Streaming → ML

Let me walk you through a journey:

Phase 1: Reporting Store

Build a data warehouse or lakehouse. Ingest key business systems (CRM, ERP, contact centre data). Transform and serve basic dashboards (how many products, how much spend, where are we today).

Value: decision boundary-based actions; time saved.

Phase 2: Domain Expansion & Streaming

Add live data (fleet sensors, IoT, user-behaviour logs). Introduce real-time alerts. Business can act quicker (if latency > threshold, then reroute).

Value: faster reaction time, cost avoided.

Phase 3: Machine Learning / Advanced Analytics

Build models: churn prediction, pricing optimisation, recommendation engines. Embed decision-making logic into the platform (if predicted churn > X, trigger incentive).

Value: new opportunity, revenue gained, risk reduced.

Few companies make it to the full “data enterprise platform” stage. If you think you have, ask yourself: Are we still measuring decision outcomes? Are people using it daily? Are we still leaning on dashboards or spreadsheets?

How to Talk About Data Engineering ROI and Value to Executives

When you’re presenting to the board or senior leadership about the data programme, use their language: time, money, opportunity. Avoid tech-speak. Frame it in business outcomes.

Example one: “By centralising data into a single platform, we’re reducing reporting cycle time by 3 days, which frees 120 person-hours per quarter, equivalent to £X in cost savings.”

Example two: “By implementing real-time pricing analytics, we expect an uplift of Y% on margin, which translates to £Z additional annual revenue.”

Use frameworks like the ones from industry (ROI framework, ROI pyramid) to back-up your case. Learn more by reading: How to Deliver a Successful Data Strategy Presentation to the Board.

Three Big Takeaways from Data Decoded

So, wrapping up the hattrick of my data decoded articles and you should now understand that:

  1. Decisions are everything. The only reason you need data is to make better decisions. Without decision-boundaries, you cannot declare you are data-driven.
  2. Platform isn’t value. Building data infrastructure is necessary, but on its own it doesn’t yield ROI. You need business usage, adoption, measurable impact.
  3. Value needs tracking relentlessly. Always link your data work to time saved, money made, or opportunity unlocked. Use simple frameworks, speak the business language, and secure buy-in.

Watch me talk about how to determine the value in a data platform here: Value of Data Platforms. And if you missed the first two articles in this Data Decoded series, catch up below.

Oakland’s Advice for Starting A New Data Programme

  • Hold a workshop to list key decisions your organisation makes every day. For each decision, ask: what data, what threshold, what action?
  • Build your platform in phases: ruling out the “acquire everything” approach.
  • Engage business users early. Set metrics for adoption, usage, decision-impact.
  • Celebrate wins, and communicate them. E.g. dashboards that work, models that deliver, cost savings, revenue gains.

And if you’re consulting or leading this work:

  • Ensure every build has a why.
  • Be the voice of value, not just tech.
  • Don’t fall into the trap of “we built it, they’ll come”.
  • Measure, iterate, expand.

The Final Word

In the world of data, everyone wants to talk about “big data platforms”, “machine learning”, “AI”, “data lakes”, “lakehouses”, “real-time streaming”. But none of that matters if what you build doesn’t change decisions or help you overcome data challenges. Because at the end of the day, you make decisions. Your organisation makes decisions. Your business lives or dies by them. Being data-driven means replacing guesswork with evidence, replacing intuition with insight and decision boundaries, and doing so consistently.

So when you hear the phrase “Data Decoded”, think less about the dazzling technology and more about the decoded decision. Think: what decision are you enabling today, with the data platform you have or will build? How will you know you’ve improved that decision? What value will result – in time saved, money made or opportunities unlocked? And when you bring in consultants or build your data team, make sure they understand: the platform is the enabler, the business decision is the destination.

That’s how you:

  • Turn “we have a data platform” into “we are a data-driven business”
  • See ROI in data programmes
  • Get value from a data platform

And that’s how data consulting really pays off.

To speak to me further about any of the challenges we’ve touched on here – or for anything else to do with data platforms – please get in touch.

The post Common Data Challenges & How to Avoid Them appeared first on Oakland.

]]>
The Four Pillars of a Data Platform https://weareoakland.com/blog/four-pillars-of-a-data-platform/ Thu, 30 Oct 2025 08:28:20 +0000 https://weareoakland.com/?p=9781 Data Decoded: How to go from “we want to be data-driven” to “we are data-driven”. We hate to break it to you, but there isn’t any natural ROI in building a data platform. The ROI happens when decisions get made from that platform. As a data engineer or data platform consultant, you build around four...

The post The Four Pillars of a Data Platform appeared first on Oakland.

]]>
Data Decoded: How to go from “we want to be data-driven” to “we are data-driven”.

We hate to break it to you, but there isn’t any natural ROI in building a data platform. The ROI happens when decisions get made from that platform. As a data engineer or data platform consultant, you build around four key functions: Ingesting, storing, transforming, and serving the data. If your platform does these four things, you’re in business. But just because you can tick all four pillars of a data platform doesn’t mean you can tick one that says it provides ROI, too.

Here’s why.

The Four Pillars of a Data Platform Build

 We’ll start with the four functions in more detail:

  1. Ingest – bringing data in from source systems (ERP, CRM, cloud apps, IoT, etc).
  2. Store – persisting the data in a suitable repository (data lake, warehouse, lakehouse).
  3. Transform modeling, cleansing, structuring, and preparing the data for use.
  4. Serve – providing access: dashboards, BI tools, ML models, APIs.

Many data engineering projects start with the diagram ingest → store → transform → serve. And that’s fine. But companies fall into the trap of believing: “we have dashboards, we have a lakehouse, we have streaming ingestion”, therefore “we’re data-driven.” But that’s wrong. If you haven’t tied it to decision boundaries and measurable outcomes, you are spinning wheels.

A Platform Doesn’t Equal Value

Data platforms have cost: Infrastructure, licences, engineers, maintenance. Without real usage and decision-making, you could have a negative ROI. For example, a data platform sitting idle still costs you £5-10k/month in cloud costs.

Jump into this further by reading: How to Manage Spiralling Cloud Costs.

What Does A Data Consulting Company Do?

A good data consultant doesn’t just build a stack. They ask: 

  • What decisions will this support? 
  • What value will it deliver? 
  • How will we know? 

A data consultant helps you avoid the “acquisitive” trap and guide you to become value-driven.

Here’s what our data consultants often advise:

Define business outcomes first

Understand what business decisions you want from the data (pricing decisions, logistics optimisation, customer segmentation, risk reduction).

Set decision boundaries

For each key metric, define the threshold to drive action.

Align data product scope to value

Don’t build everything. Prioritise data products (dashboards, ML models, APIs) that map to time/money/opportunity.

Ensure adoption & governance

A platform is only useful if business users engage with it.

Track ROI continuously

Measure before-and-after, quantify value, build case studies. For example, when consultants deliver dashboards & platforms they embed tracking: “We reduced reporting cycle time by X hours, we reduced cost of chasing data by Y, we improved conversion by Z%.” 

Good consulting keeps the focus on action, not just architecture. Learn more: What is Data Platform Architecture?

How Do You Get Value from a Data Platform?

Let’s break down meeting the four pillars of a data platform into actionable steps:

Step 1: Start with Decisions

You must map which decisions your organisation needs to make routinely. For example:

In utilities: “If usage anomaly > 5%, then investigate the customer meter.”

In transportation: “If fleet-cost per km > X, then re-route fleet or renegotiate contract.”

In tech: “If process latency > X, trigger optimisation run.”

From there, you define the data that will inform that decision.

Step 2: Build Decision Boundaries

As noted earlier, define thresholds in advance. Know what numbers trigger what actions. This ensures your dashboards don’t just show “what happened”, they support “what we will do”.

Step 3: Build the Platform Lean

Instead of “ingest everything”, you prioritise the data sources that serve those decisions. Map ingestion, storage, transform, and serve accordingly. Then build the initial version – it may just be a reporting store – then grow from there: more domains, real-time streaming, ML models.

Step 4: Operationalise the Platform

You need adoption. Business users must use the dashboards, models must deliver insights, the thresholds must be monitored and acted upon. If your data platform is just “nice to have”, it won’t deliver ROI. Remember: value comes when decisions change behaviour. 

Step 5: Measure Value

You must link back to time, money or opportunity. Use frameworks to estimate value – for example:

Data ROI = (value from data initiative – cost of initiative) / cost of initiative

Track metrics such as: time saved, cost avoided, revenue gained, risk reduced. Get evidence around before / after.

Step 6: Iterate and Expand

Your first data product may be reporting. You then add real-time streaming, machine learning, and automated decisioning. The “lighthouse model” (start small, show value, expand) works better than big-bang POCs that die.

Two down, one to go! For more insight like this, please read the third and final blog in this data decoded series: Common Data Pitfalls & How to Avoid Them. You can also send your questions to me or speak with the rest of the Oakland consulting team by getting in touch – we’d love to hear from you.

The post The Four Pillars of a Data Platform appeared first on Oakland.

]]>
Data Decoded 2025: How to Get the Most from Data Consulting https://weareoakland.com/blog/get-the-most-from-data-consulting/ Thu, 30 Oct 2025 08:27:14 +0000 https://weareoakland.com/?p=9777 How to see ROI, extract value from a data platform and get more from your data consultancy, as explored by our Principal Engineer Advocate, MLG, at Data Decoded 2025. In today’s world, you don’t run a business unless you make decisions. It doesn’t matter what industry you’re in – utilities, transportation, computing, retail – you...

The post Data Decoded 2025: How to Get the Most from Data Consulting appeared first on Oakland.

]]>
How to see ROI, extract value from a data platform and get more from your data consultancy, as explored by our Principal Engineer Advocate, MLG, at Data Decoded 2025.

In today’s world, you don’t run a business unless you make decisions. It doesn’t matter what industry you’re in – utilities, transportation, computing, retail – you are in the business of making decisions. 

Whether you’re setting unit-rates of electricity in utilities, organising a fleet in transportation, or optimising a process in the tech/computer industry, the truth is that you are making decisions. And to compete and stay relevant, you try to be data-driven. Because philosophically that’s the only way humans really know how to operate, right?

The Purpose of Being Data-Driven

When I say “data-driven as a business”, I mean this: for many of the decisions we make as a company, as an employer, as a person, we stop trusting “vibes” or our “spidey sense” about a customer’s behaviour, or whether some process will succeed or fail. Instead, we lean on evidence. We use the data. We say, “This person matched profile A rather than B. Therefore, we think this person will buy A rather than B because we have evidence and data to support that.”

You could spend a lifetime writing books on the subtlety of evidence-driven vs bias-driven decisions. But we don’t have the luxury of endless pages. So, let’s unpack what it means to truly become a data-driven organisation.

Clear Decision Boundaries

First, if you’re going to drive value from data, you need to adopt some ground-rules. In particular, you need clear decision boundaries. 

I want you to reflect on the worst dashboards you’ve seen: the chart-junk, the yellow warning‐lights, the numbers splattered without interpretation. Those dashboards are not enabling decision-making. They’re confusing. 

If you’re going to be evidence-driven, you must know in advance:

  • If this number > X, we’ll do Action A.
  • If this number < X, we’ll do Action B.

At platform, engineering and analytics levels, this means knowing your decisions before you build the data asset. Because when you build without knowing the outcome, you’ll create a data platform, you’ll move numbers, you’ll host dashboards…

…but you won’t get value.

Data Decoded: Stop Acquisitive, Start Inquisitive

One of the biggest pitfalls in data programmes is what I call an ‘acquisitive data strategy’. That is: “Let’s take every dataset, grab every number, throw it into a warehouse, and then maybe something good will happen.” 

No. 

You need an inquisitive strategy. Ask yourself:

  • What decisions will you support? 
  • What thresholds matter? 
  • What outcomes will change? 

Then drive your data platform accordingly.

When businesses implement a data platform, they often think: “Here’s our cloud budget, build everything, feed everything into the data warehouse, hook up dashboards, we’ll iterate.” But a platform alone doesn’t give ROI. Tools alone don’t deliver value. Without decision-boundaries, without usage, without an evidence base, you’re just moving data around from one expensive database to another expensive database.

What Does “Value” Really Mean?

So, let’s get into value. You may have heard the initialism “ROI” (return on investment). In the data world, ROI is sometimes misused or misunderstood. Value, when you’re building data platforms or data products, comes in (at least) three flavours: time, money and opportunity. Yes, they’re all about money if you dig into them – but they come at you in different ways.

Time saved 

If your data platform or analytics process saves hours, days, weeks of manual reporting or intervention, that frees people to do higher-value work.

Money made

If you can use your data to optimise pricing, reduce cost of goods sold, improve marketing performance, increase conversions, then you’re generating revenue (or preventing loss).

Opportunity unlocked

Sometimes you build a platform that enables new products, new customer segments, new business models – and all of these are valuable too, even if it’s future-facing.

If you’re working as a data engineer, a data platform consultant, or leading a data programme, you must ask: Which one of these am I addressing when building this feature, this platform, this dashboard? If you cannot answer that, you should reconsider your life choices…

…(well, maybe!).

For more insight on the value of data, please read: Why is Data Important for Business?

What We Do Day-to-Day

At the end of the day, what you and I do is we make decisions. Whether you’re in utilities, transportation, tech, retail, you decide. And your ambition is to decide based on evidence, not guesswork. When you embrace data, you move away from trusting your “gut” to saying: “Here’s the evidence, let’s act.”

To become data-driven means that for many of the decisions now, as a company, we’re going to stop using vibes, stop relying on “this customer might buy” because our spidey sense tells us so. Instead we say: “Look at the data. Based on this profile, we believe they will buy A, not B.” Because we have the evidence.

That’s the first of my Data Decoded series finished, but there’s plenty more like this in the second blog: The Four Pillars of a Data Platform. You can also send your questions to me or speak with the rest of the Oakland consulting team by getting in touch – we’d love to hear from you.

The post Data Decoded 2025: How to Get the Most from Data Consulting appeared first on Oakland.

]]>
The Role of Databricks Architecture in Data Engineering https://weareoakland.com/blog/advanced-data-engineering-with-databricks/ Mon, 04 Aug 2025 09:43:33 +0000 https://weareoakland.com/?p=9665 Since it was founded in 2013, Databricks has revolutionised enterprise data management and analytics. Built on Delta Lake, an open source storage format, the set of data engineering tools it prides itself on processing enormous amounts of data, then transforming them into datasets that are primed for exploration via machine learning (ML) models. At Oakland,...

The post The Role of Databricks Architecture in Data Engineering appeared first on Oakland.

]]>
Since it was founded in 2013, Databricks has revolutionised enterprise data management and analytics. Built on Delta Lake, an open source storage format, the set of data engineering tools it prides itself on processing enormous amounts of data, then transforming them into datasets that are primed for exploration via machine learning (ML) models.

At Oakland, anything and everything data is at the core of the data and AI consultancy services we provide. When it comes to data engineering, Databricks is a central block in how we build advanced data platforms to provide actionable data insights for clients. 

In this article, we dive into the details of building a data platform with Databricks, including:

  • The role of Databricks in data engineering
  • Why Databricks and ETL (extract, transform, load) are a match made in data heaven
  • An overview of Azure Databricks
  • The pros and cons of Databricks
  • Use cases of Databricks

Let’s start.

The Role of Databricks in Data Engineering

Databricks plays a central role in modern data engineering by providing a scalable, high-performance platform. Built on Apache Spark (which is around ten times faster than traditional SQL databases), it enables teams to ingest, transform, and process vast volumes of structured and unstructured data efficiently. 

With features like Delta Lake for reliable data storage, and Photon for accelerated SQL performance, Databricks powers robust ETL (extract, transform, load) pipelines, real-time processing, and advanced analytics. Its unified workspace supports collaboration across data engineers, analysts, and data scientists, making it a key component of enterprise data platforms.

Databricks and ETL: A Match Made in Data Heaven

As a cloud-based platform, Databricks lends itself to ETL workflows – in fact, several of its tools and features have been specially designed with ETL pipelines in mind. So if slicker data extraction, transformation, and loading is important to your data activities, a platform engineered using Databricks could be a perfect fit. 

Some of the benefits (and the features that enable them) are listed below:

Easier ETL development: 

Thanks to Databricks Lakeflow Declarative Pipelines (previously Delta Live Tables), the operational complexities of ETL processes are automated. You define what should happen, not how, reducing boilerplate code and enabling ETL in SQL or PySpark, speeding up development cycles and reducing operational overhead. ETL can be written in Spark or SQL, too, for extra flexibility.

Streamlined workflows

ETL tasks, analytics, and machine learning pipelines are all orchestrated in Databricks Lakeflow Jobs (previously known as Databricks Workflows).

More focus on data quality

Thanks to features like Lakeflow Declarative Pipelines and automated data quality (DQ) testing, Databricks reduces the need for engineers to manage pipeline infrastructure or check DQ, freeing up time to deliver high-quality data.

What is Azure Databricks?

Given its power, it was only a matter of time before Microsoft jumped on the Databricks capability. In 2017, they became a first-party provider of Databricks’s cloud-base platform, integrating it with its own Azure cloud services. The result? Azure Databricks, the open analytics platform that allows you to build, deploy, share, and maintain enterprise-grade data, analytics, and AI solutions at scale. 

Naturally, as a Microsoft Partner awarded the Analytics on Microsoft Azure specialisation, we were super excited about the integration! Azure Databricks is another building block in our data engineering toolkit, allowing us to engineer data platforms at enterprise scale. Not to mention the immense potential it’s opening up for our customers and their data assets.

Our blog, ‘How to create a secure Azure data platform’, looks at Azure services in more detail.

“Our collaboration with Microsoft builds on our momentum as a leading cloud platform for Apache Spark-based analytics. The ability to provide our Unified Analytics Platform to all Microsoft Azure users in such an integrated fashion is invaluable to end users looking to simplify big data and AI.” 

Ali Ghodsi, Co-Founder and CEO of Databricks

Databricks Use Cases

With this in mind, let’s cut to three of our recent use cases using Databricks as part of our advanced data engineering service.

1. Building a sustainable, long-term data platform for Network Rail

Data platform engineering is an investment, so you need to make sure the technology is set up for future success. Databricks enables an open approach, reducing the complex nature of being ‘locked in’ that comes from using a more traditional platform vendor. Something our client, Network Rail, knew all too well.

Like many other large organisations with legacy data platforms, Network Rail was struggling to access data, which made extending the capabilities of their datasets difficult. Using Databricks, we built an open data platform architecture for the rail services provider, which has:

  • Automated manual processes, freeing up valuable time and resources to spend elsewhere
  • Reduced lead times by merging two loosely data pipelines into one
  • Enabled smarter decision-making, thanks to the deployment of advanced analytics and ML tools

Yet that’s just the start – Click here to read the full case study.

2. Informing sales strategies for a leading provider of IT infrastructure

Sales team struggling to extract data from multiple sources? We recognise the challenge, and it’s one we helped a leading provider of IT infrastructure overcome. 

After we designed the IT service firm’s new data analytics platform, we leveraged the Databricks stack to build a machine learning and data science model. Their sales team now have access to far richer insights, driving better margins for the overall business. These insights include:

  • A customer’s tendency to buy certain product categories 
  • A highlighting system of the products that customers are likely to buy
  • Product penetration and available spend information, so staff can quickly spot where to focus time and energy

3. Driving an ROI increase of £150m+ for Yorkshire Water

As part of an overall data transformation programme, we developed a new data platform for Yorkshire Water. Databricks was primed to be the enterprise data architecture for the utilities company and was pivotal to the design and build of their new, strategic data platform. 

In total, ROI from the overall business transformation has exceeded £150m.

What are the Pros and Cons of Databricks?

It’s fair to say our Databricks and Azure Databricks use-cases and results speak for themselves. However, it’s important to weigh up the pros and cons of any data architecture to make sure you’re choosing the best fit for your business needs. We’ve outlined some of the major pros and cons of Databricks below to give you a better understanding of whether it’s right for you or not.

Pros

Advanced data governance capabilities, such as data lineage, roles, and permissions thanks to an in-built Unity Catalogue

Ease of scaling and maintenance

One unified platform for batch, streaming, ML, AI, and analytics

Native integration with all major cloud platforms and PaaS, plus native DevOps and Git support

Eliminates data silos by using Data Lakehouse architecture

Provides a collaborative approach to Data Warehousing in a database

The ETL process is in-built and in one place, omitting the need for another tool

Features are continuously updated and added 


Cons

Cost, especially at scale

Higher barrier to entry for non-developers

More suitable for bigger datasets

Constantly evolving product, so you need to have the time and resources to dedicate to understanding these changes

Data Engineering Advice

Of course, for more advice on Databricks and data engineering, or to speak to us about your needs for a data platform, please get in touch with our friendly team. That’s what makes us Oakland, everything data.

The post The Role of Databricks Architecture in Data Engineering appeared first on Oakland.

]]>
What is Data Platform Architecture? https://weareoakland.com/blog/what-is-data-platform-architecture/ Fri, 06 Jun 2025 09:07:20 +0000 https://weareoakland.com/?p=9566 What is Data Platform Architecture? Is your data architecture designed to ensure scalability, reliability, and long-term success? We know that organisations are increasingly investing in data platforms to bring their data together in a world driven by data. And while this enables you to gain the insights you need to make better business decisions based...

The post What is Data Platform Architecture? appeared first on Oakland.

]]>
What is Data Platform Architecture?

Is your data architecture designed to ensure scalability, reliability, and long-term success? We know that organisations are increasingly investing in data platforms to bring their data together in a world driven by data. And while this enables you to gain the insights you need to make better business decisions based on knowledge (not assumption), building a data platform isn’t just about choosing the latest tools or storing massive amounts of data. 

It’s also about choosing data architecture that’s most relevant and applicable to your business. 

If the architecture doesn’t fit in with your overall enterprise, the platform investment risks ending up as a missed opportunity. 

Read on for everything you need to know about data platform architecture, from why it’s important when building a data platform to key considerations when choosing the right architecture for your business.

What is Data Architecture?

Data architecture is the strategic design and structure of how data is collected, stored, managed, and used within an organisation. It acts as a blueprint  for how data flows through systems and how different components, like databases, data warehouses, data lakes, and analytics tools, interact with each other. Think of it like constructing a building: without a solid architectural plan, you risk creating something unstable, inefficient, or impossible to live in. 

Read our article to learn more about building or buying a data platform.

Why is Data Architecture Important?

At the very heart of a data platform is a core set of capabilities (ingestion, storage, processing etc). However, understanding your business needs and how these capabilities fit within the wider enterprise is fundamental to success of the platform. 

For example, think about the source systems where your data is held. If 90% of your data is held within a core Enterprise Resource Planning (ERP) solution (such as SAP), there is limited benefit in moving all your data to a separate data solution – especially when you may be driving value from the wider SAP estate, such as SAP Analytics cloud. 

Likewise, if most of your use cases require real-time data to support operational decision making, understanding where your data is and how you’d like to use it will immediately shape the platform design. Knowing your business, its data and application estate, and case for change are foundational pieces that must be used to define the architecture of the platform, rather than trying to force the business to adopt something that isn’t fit for purpose.  To continue our house analogy, there is no point having an amazing surround sound 4k home cinema if there are no plug sockets nearby to plug it in!

Scalability

As data volumes grow, your platform must scale without breaking. A well-architected platform anticipates growth and supports horizontal scaling, distributed processing, and modular components that can evolve independently. Understand what else you need to consider when designing and building your Data Platform:

Performance

Poor data architecture can lead to bottlenecks, slow queries, and frustrated users. Optimising data flow, storage formats, and compute resources ensures your platform performs well under pressure.

Data Quality and Consistency

Investing in architecture and setting out your data architecture roadmap can help enforce data governance and standardisation, not to mention data quality. By defining clear data models, validation rules, and lineage tracking, you ensure that data remains accurate, consistent, and trustworthy. This trust is critical if people are going to use the platform, otherwise, it runs the risk of being a very expensive storage solution. 

Security and Compliance

With increasing regulations, like GDPR, and the increased cyber threat, data security is non-negotiable.  It seems like we’re seeing organisations, such as Marks & Spencer and the Co-op, in the press due to a data breach or cyber attack. A strong architecture can help minimise the risk of this.  Considerations in this area includes access controls, encryption, auditing, networking  and compliance mechanisms from the ground up.

Beyond this, the architecture also needs to think about how these policies & processes work with wider enterprise governance requirements. For example, your platform needs to be able to facilitate a ‘right to erasure’ request, even though it is not the source system that is generating the data.

Flexibility and Integration

Modern data platforms must integrate with a variety of tools, such as BI dashboards, machine learning models, APIs, and more. Focus on data access both to and from the platform, and do not see the data platform as the end point, merely part of the wider data ecosystem.

Cost Efficiency

Without thoughtful architecture, cloud costs can spiral out of control. In particular, where organisations are moving from on-prem to cloud, this shift from CAPEX to OPEX needs careful consideration. While running your platform, there is a constant need to ensure that you are optimising your services. Efficient data partitioning, storage tiering, and compute resource management help keep costs predictable and manageable. 

Find out the data architecture we used to build a platform for Income Analytics: Creating catalysts for growth with Income Analytics.

Key Components of a Good Data Platform Architecture

  • Data Ingestion Layer: Handles batch and real-time data collection from various sources within your organisation. 
  • Storage Layer: Supports structured, semi-structured, and unstructured data. 
  • Processing Layer: Enables data transformation, enrichment, and aggregation.  Effectively adds business meaning to your data. 
  • Serving Layer: Powers analytics, dashboards, and data science workloads and enables users to drive benefits. Ultimately, if users aren’t accessing the data, you have created a very expensive garage with boxes full of old magazines and Christmas decorations. 
  • Governance Layer: Manages metadata, lineage, access control, and compliance, all key capabilities of a data platform.

Final Thoughts on Data Platform Architecture

A well-designed architecture is the foundation of a successful data platform. It’s not just a technical necessity – it’s a strategic advantage. By investing in thoughtful architecture from the start, you ensure that the solution you deploy is both fit for your organisation, and provides users with the capabilities to drive insights and make better decisions.

Contact us to find out how we can advise you on the right data architecture for your platform or check out our guide on building one yourself!

The post What is Data Platform Architecture? appeared first on Oakland.

]]>
Should you Adopt Microsoft Fabric as your Data Platform? https://weareoakland.com/blog/adopt-microsoft-fabric-as-your-data-platform/ Tue, 07 Jan 2025 14:29:08 +0000 https://weareoakland.com/?p=9239 As Microsoft Partners; here at Oakland we’re being asked by more and more clients if Microsoft Fabric could help them to drive better more profitable outcomes from their data? Source: https://www.microsoft.com/en-us/microsoft-fabric Microsoft Fabric has been out on general release for a while now and is increasingly seen as a serious alternative to a bespoke Data...

The post Should you Adopt Microsoft Fabric as your Data Platform? appeared first on Oakland.

]]>
As Microsoft Partners; here at Oakland we’re being asked by more and more clients if Microsoft Fabric could help them to drive better more profitable outcomes from their data?

Source: https://www.microsoft.com/en-us/microsoft-fabric

Microsoft Fabric has been out on general release for a while now and is increasingly seen as a serious alternative to a bespoke Data Platform. Promising the capabilities of a one stop shop for all your data needs, from simple dashboards to advanced data science workloads. But is it right for you? Even if you currently use the Microsoft data stack (Synapse, on-premises SQL Server) the answer isn’t an automatic yes, as we’ll explain below…

What the Problems Does Microsoft Fabric Solves?

For existing Microsoft customers, Fabric is certainly being touted as a no-brainer. Integrating with Microsoft 365, enabling collaboration through Teams, Copilot and SharePoint users can quickly access insights directly using familiar tools. 

Managing multiple tools for data visualisation, governance, and collaboration can be cumbersome and costly therefore bringing them all under the Microsoft umbrella can be seen as easier and more cost effective.

The popular Power BI integration can provide visualisations and reporting tools directly tied into the platform.

And let’s not forget Data Governance – Fabric promises robust data governance, ensuring compliance, security, and data lineage across all processes.

  • Built-in Security: Unified security measures like role-based access and encryption.
  • Unlike several alternative data platforms, there is no extra cost or much additional configuration for Microsoft Entra/Office 365 Single Sign-On (SSO).
  • Data Lineage: Comprehensive tracking of data transformations and movement.
  • Purview Integration: Facilitates data cataloguing and compliance tracking.

Fabric uses OneLake, a universal data lake that centralises data storage and enables cross-platform access without creating silos. This streamlines data integration from multiple sources, including on-premises, cloud, and SaaS systems.

Many organisations we see struggle with siloed data across on-premises, cloud, and SaaS platforms, making integration complex and time-consuming. A traditional data platform can be the answer but often by the time the platform is created the technology is out of date or the people and processes needed to use the platform to its full potential are disengaged and doing their own thing.

Direct Lake Implementation, while having a few limitations (largely it’s not great for reasonably complex Power BI dashboards at present) does offer near real-time loading of large data sets with little to no configuration, which can be a major productivity boost. How often have your data teams been left twiddling their thumbs waiting minutes or even hours for your Power BI datasets to refresh because of changes made to your source data with direct import?

And as 2025 is the year of Agentic AI we can’t fail to mention the challenge of incorporating AI into existing workflows. Of course, Microsoft Fabric has a solution for that with its AI Co-pilot integration, while only currently available at F64 capacity / price tier (£8k a month) this gives us a tantalising glimpse into major productivity boosts for your engineers and analysts. Sound good so far? These are also improvements over Microsoft’s other data platform, Azure Synapse Analytics, but there is also improved Spark performance, more integrated real-time analytics, Lakehouse shortcuts for easier loading of data in the Data Lake, plus many more features and by the time this blog is published no doubt Microsoft will have added more new features to the mix.

How Hard is it to Implement Fabric?

Getting started super easy: deploying a capacity and workspace only requires a few clicks.

Although we’ve noticed implementing a typical development, test, and production software lifecycle in Fabric can be tricky as git integration and/or pipeline deployments are either in preview or straight up missing for some Fabric features.

Source: https://learn.microsoft.com/en-us/fabric/cicd/git-integration/intro-to-git-integration?tabs=azure-devops#supported-items

You also must be careful how you implement Fabric; it has many options! For example, there are three different ways of doing data transformations:

  • Python Notebooks
  • SQL Queries
  • Dataflow Gen 2 (low-/no-code)

Each has their own use cases, personas, pros and cons, so we recommend you understand each of them, including trying them out before committing to one (or more). This goes for many other aspects of Fabric (ingesting data, CI/CD process, data science workflow, etc.).

Another thing to note is that if you want private endpoints or on-premises data, they are more work to implement and can be tricky to do, especially if you’ve not implemented them before.

What does Fabric Cost?

Unlike Synapse, Databricks and Snowflake which offer pay-as-you-use/pay-as-you-go with auto-shutdown for compute costs, Fabric compute capacity charges are like a cloud Virtual Machine (VM) or a Netflix subscription: it’s always on and racking up costs unless you pause it, which turns off all compute and most of its features.

Just like a cloud VM, they are quick and easy to build, start, pause, and destroy via the Azure portal and can build more than one of them in an Azure tenant, so we find them great for testing out ideas.

Source: https://learn.microsoft.com/en-us/azure/architecture/analytics/architecture/fabric-deployment-patterns#pattern-3-multiple-workspaces-backed-by-separate-capacities

This does mean, however, production Fabric capacities are often left on 24/7, which means more simple and predictable cost forecasts but potentially not as efficient if you’re not using all or most of your capacity 24/7 as serverless pay-as-you-use, especially as serverless Databricks and Snowflake are generally cheaper per hour than their non-serverless workloads.

Common Data Platform compute costs throughout a typical (Mon-Fri) working day.

To be fair, if you’re running Fabric 24/7, you can get 41% discount for a year’s reservation and you could scale & pause compute on a schedule with Fabric’s REST API, especially for development workloads. Though doing this starts to impact Fabric’s main selling point – it’s easy to run and maintain.

UK South Capacity Pricing. Source: https://azure.microsoft.com/en-us/pricing/details/microsoft-fabric/

A single compute capacity can also cause problems as you scale in users and workloads, as workload management is a bit limited in Fabric right now, and your capacity compute powers almost everything. This can lead to trying to figure out who or what has eaten all the compute. This can also lead to further costing inefficiencies as you scale up pricing tier /capacity or cloning workspaces to another capacity avoid having your work blocked by someone else at busy periods.

There are other potential Fabric costs to be aware of:

  • OneLake storage costs, which are similar or slightly higher (-0 to 20%) than Azure Data Lake costs.
  • You’ll still need Power BI Pro licences for your Power BI developers.
  • Gateway costs if you require a vNet/on-premises gateway.
  • Private Endpoint costs, if required.

The above costs are often small compared to capacity costs, with capacity costs often being approx. 70% to 95% of your total run costs for Fabric depending on the options you choose.

Also note, many preview features are locked behind F64 capacity / pricing tier: we’ve seen some clients opt for this capacity to get the additional features, even though they don’t need the compute.

Who is Fabric Suitable For?

Existing Fabric/Power BI users who want to expand their data capability without going through the extra hassle of adopting another tool.

Migrating from low-/no-code focused products then Atleryx and Talend are alternative options, especially if you’re looking for a data platform with more Azure integration.

Migrating from an on-premises or cloud SQL Server will give you more data engineering, science and governance capability compared to on-premises tooling like SSIS.

You might also think existing Synapse users are great for Fabric since they are made by the same company, and Fabric is in many ways an evolution of the Synapse features previously mentioned. This is very true: a migration to Fabric will be easier than to any other product. Though beware, as mentioned above the cost model is different in Fabric, and there are some features in Synapse not or heavily changed in Fabric* that could block a migration or require significant rework (for example a lack of parameters in pipeline connections).

We don’t believe existing or target (large enterprises) Databricks or Snowflake customers will be adopting Fabric immediately: while the foundations are there, I think there are too many essential features in preview or missing, with examples such as git integration and workload management, which I mentioned above.  

But and there is always a but the only caveat we’d add is if you have a business need for using Direct Lake, we can see some large companies using Fabric with Databricks/Snowflake in tandem – Microsoft even makes this easy by having great integration with Databricks and Snowflake.

Source: https://learn.microsoft.com/en-us/fabric/onelake/onelake-overview

*As of January 2025, we suspect this sentence will need changing in the coming months as Microsoft is working hard for feature parity with Synapse.

Summary

While Fabric is a step forward from Synapse, there are differences you’ll have to think about before migrating, and it doesn’t make Azure Databricks redundant either, especially for large, complex enterprises.

And while we’ve found it a great experience to get started with Fabric, there can be a few challenges and decisions to be made before reaching production.

If you would like a deeper dive into comparison of Fabric against its competitors and advice on how best to implement Fabric, feel free to reach out or download our Data Platform guide.

Interested in unifying your data estate with Microsoft Fabric? Join our workshop to discover how Microsoft Fabric can streamline your data integration and management. We’ll explore its features and assess whether it’s the ideal solution for your organisation’s needs.

The post Should you Adopt Microsoft Fabric as your Data Platform? appeared first on Oakland.

]]>
Should you use Data Lakehouse instead of a Data Warehouse and / or Data Lake?  https://weareoakland.com/blog/should-you-use-data-lakehouse-instead-of-a-data-warehouse-and-or-data-lake/ https://weareoakland.com/blog/should-you-use-data-lakehouse-instead-of-a-data-warehouse-and-or-data-lake/#respond Fri, 08 Mar 2024 14:39:42 +0000 https://www.theoaklandgroup.co.uk/?p=7265 Should you use Data Lakehouse instead of a Data Warehouse and / or Data Lake?  Intro When using your Data Platform to improve your Business Intelligence with useful dashboards and reports, you’ll more than likely want to use a Data Warehouse. Add on your data science builds, and storing your raw data cheaply, plus adding...

The post Should you use Data Lakehouse instead of a Data Warehouse and / or Data Lake?  appeared first on Oakland.

]]>
Should you use Data Lakehouse instead of a Data Warehouse and / or Data Lake? 

Intro

When using your Data Platform to improve your Business Intelligence with useful dashboards and reports, you’ll more than likely want to use a Data Warehouse. Add on your data science builds, and storing your raw data cheaply, plus adding a Data Lake just for good measure, and the costs soon start to add up. Running both in tandem on a Data Platform can have serious costs and maintenance associated.

So, can you have the best of both worlds with the Data Lakehouse? And what is the best Lakehouse to use

Before we answer those questions, we must ask “What is a Data Warehouse, Data Lake and a Data Lakehouse?”

What is a Data Warehouse, Data Lake and a Data Lakehouse?

Data Warehouse is a data architecture that has been around since the 90s and is still relevant today. It is a means to store tabular data so it can be easily used by business intelligence applications such as Tableau or Power BI, web applications, and even other data warehouses. The three most common Data Warehouse architectures are Kimball Star Schema, Data Vault and One Big Table.

The name is also confusingly used to identify a type of Database, such as AWS Redshift, Azure Synapse and Snowflake, which specialise in storing and querying large amounts of data.

Data Warehouses have their issues; they can be more expensive than a Data Lake when processing large amounts of data, and work best when data is of reasonable quality and in a tabular structure.

Architecture of a simple Data Platform using just a Data Warehouse. 

So, along came the Data Lake to help ease these common pain points:

  • Data Scientists needing to be able to process large amounts of raw data of dubious quality. 
  • Increasing requirements for storage of non-tabular data sources. 
  • The need for data storage that is more flexible in structure and schema. 
  • The need to store data that might be needed at later date, for example for auditing, but have a low setup and maintenance cost (little or no ETL process needed compared to a Database). 

Data Lake is just a distributed file system at its heart, usually hosted in the cloud in AWS S3 or Azure Data Lake, with large files split by a key, so you can save on processing costs by loading the partitions you need.

Data Lakes also generally have more flexibility in that it can store an unlimited number of file formats and offer a common interface to its storage that allows you to use many compute engines. This often called separating storage from compute, which has become so popular that many Data Warehouses offer this too now. Data Lakes can also easily store non-tabular data (images, videos and music) that Data Warehouses cannot without some pre-processing.

However, without Delta Lake it cannot easily or efficiently do row level updates and inserts, nor connect easily to business intelligence applications, that a Data Warehouse or Database can do.

Architecture of an example Data Platform using both a Data Lake and Data Warehouse.

What is a Data Lakehouse?
A Data Lakehouse is an open data management architecture that combines the flexibility, cost-efficiency, and scale of Data Lakes with the data management and ACID transactions of Data Warehouses, enabling business intelligence (BI) and machine learning (ML) on all data.  

What is Databricks Lakehouse? 
Until a few years ago, Databricks was mainly designed as an easy way to run Spark, a distributed data processing library for large scale Data Engineering and Data Science. It worked mainly in tandem with a Data Lake, with similar advantages and drawbacks. 

In 2019 Databricks released Delta Lake, a file format with attributes only found previously in Databases and Data Warehouses as mentioned above. Combined with Spark to process and transform a wide variety of data, this gave birth to the Data Lakehouse. 

Today, Databricks has a fully featured SQL Data Warehouse, enterprise security, data governance with Unity Catalogmany data connectors, as well as the ability to output data to Power BI and Tableau, so it can meet all common data use cases. 

Architecture of an example Databricks “Lakehouse” using Spark as the processing engine and Delta Lake as storage. 

For those looking at building a Data Mesh, Databricks has federated query in preview, though Delta Lake also has connectors for TrinoStarburst and Dremio so you can join up many Data Lakes across your organisation: 

Architecture of many Lakehouse Data Products in a Data Mesh – the query layer and governance layer will have access to all Data Products, limited by access permissions.  

Will I still need a Data Warehouse? 

Maybe, but note it may take some time for a data team used to Databases/Data Warehouses and SQL to convert to Data Lakehouse. Here at Oakland we feel it is still easier to set up and optimise Cloud Native Warehouses like Snowflake and Google Big Query, than Databricks, as there are fewer moving parts.  

These maintenance costs can far outweigh the benefits of the Lakehouse, generally at smaller scales and data complexity. 

Also, while we’ve seen first-hand that Lakehouse can be the cheaper and more performant option than a Data Warehouse, this hasn’t been the case 100% of the time and you should do your own testing, as performance and cost heavily depends on the data you use and the environment you operate in. 

Can I build a Lakehouse somewhere other than Databricks? 

Yes, Delta Lake is open source and can be used in many different data compute products which are listed below. However, Databricks has built in special optimisations just for Databricks and a robust user interface to manage the Lakehouse. So, it is likely running Delta Lake will be slower and could be harder to maintain elsewhere. 

Example the Databricks user interface for datasets showing the schema and a sample of the dataset. 

Also note that Databricks is a general compute engine rather than a database or programming interface: it can run SQL, Pandas, Ray, Spark, most of the popular data science libraries, do graph analytics, geospatial, IoT, near-real time streaming and import almost any Python, Java, R or Scala library. Databricks’ main benefit to us is its extreme versatility, potentially reducing costs by not having to maintain separate business intelligence and data science data processing applications.    

Also, Databricks is in strong position to customise Large Learning Models (LLMs) like ChatGPT, with its general compute and strong MLflow integration, so you can pick the best open-source AI models and tune it with your organisational data in a highly efficient way using MLOps.  

However, if you’re already using one of the Lakehouse alternatives listed below, it may not be worth adding Databricks to your Data Platform. 

The alternatives to Databricks Lakehouse are:  

  • Starburst, like Databricks, is a cloud neutral and cloud native compute engine with a full suite of enterprise options and data connectors. It has Delta Lake and Iceberg connectors that can be fully controlled with a SQL API.  
  • Azure Synapse has the option to use its own Spark Engine, can import Java and Python libraries, and has Delta Lake Integration too. Has excellent integration with rest of Azure.  
  • AWS Glue allows you to use Delta Lake in S3. Has excellent integration with rest of AWS. 

Some may say Pandas or DuckDB can be a Data Lakehouse, though from our research in May 2023 they cannot do transactions or merges on a Data Lake file (Delta Lake, Iceberg, etc.) so have been excluded from the above – they still have their own use cases though.  

Summary

In short, like with other data products and architectures, the answer is it depends on the makeup of your data team, security, the size and structure of your data, and how the data is used among many other factors.  

If you are consuming a lot of data in your data platform, struggling to manage both a Data Lake and Data Warehouse at the same time, or trying to figure out how to use advanced analytics like Machine Learning with your data, Data Lakehouse is in our opinion a convincing proposition.   

We also find ourselves recommending Databricks more often than the alternatives as it offers the most complete Lakehouse solution, though competitors are quickly catching up and offering a near as good as experience as Databricks, so the choice isn’t as easy to make as it was in 2021 when we first wrote this article. 

The post Should you use Data Lakehouse instead of a Data Warehouse and / or Data Lake?  appeared first on Oakland.

]]>
https://weareoakland.com/blog/should-you-use-data-lakehouse-instead-of-a-data-warehouse-and-or-data-lake/feed/ 0
Should you build or buy your data platform? https://weareoakland.com/blog/should-you-build-or-buy-your-data-platform/ https://weareoakland.com/blog/should-you-build-or-buy-your-data-platform/#respond Wed, 08 Nov 2023 13:36:50 +0000 https://www.theoaklandgroup.co.uk/?p=7790 “Build Versus Buy” is More a Scale of Options than a Binary Decision As consultants, we’ve been involved in helping our clients to decide which is the best software to invest in (or not invest in some cases). This is quite a responsibility as we are putting our reputation in the hands of a vendor,...

The post Should you build or buy your data platform? appeared first on Oakland.

]]>
“Build Versus Buy” is More a Scale of Options than a Binary Decision

As consultants, we’ve been involved in helping our clients to decide which is the best software to invest in (or not invest in some cases). This is quite a responsibility as we are putting our reputation in the hands of a vendor, and we know that tech vendors are amazing at making claims. We see it in every trade show we attend. Therefore, it can be impossible to verify all the features of every product.

From our experience, this can be difficult to get right. So here are a few lessons in buying software (or not) to avoid making costly mistakes.

Limiting Your Potential Options is Sensible but Dangerous

One sales tactic is to present two options. This makes sense in some ways: management is time constrained and doesn’t want to assess all 100+ options that are possible, but this can lock you into a limited way of thinking, especially if you only assess two extreme “buy“ and “build” options.

Diagram to show investments in software for a data platform on a scale

Scale of Product Types from Build to Buy. By Jake Watson.

So you want to look at a variety of options, and for most organisations, look at options across the scale. We find a managed Platform as a Service (PaaS) in the middle of the scale, fits well for most but not all!

Think Beyond a Product

You may also want to consider a mix of both build and buy components; this can get you closer to fulfilling a set of complex requirements in a reasonable budget and timeframe.

One of our favourite examples is Fivetran, a vendor that offers a super quick and easy way of getting source data into your Lakehouse or Warehouse. However, it can cost a lot at scale, so you may end up doing the math and finding that building your own ingestion is more cost-effective for some, but not all, sources. Or try open source (AirbyteMeltano).

And yes, maintaining two or more systems is generally worse than one, but that should be a guideline, not a rule, as sometimes, you are the exception.

There are Too Many Choices Though!

You can go too far the other way and spend months going through dozens of options out of 1400+ products that exist in data and AI.

One example of the variety of options available to you is real time data processing, where you can choose from least to most managed:

*For those who haven’t come across IaaS, PaaS and SaaS we recommend this article.

We’ve probably missed a few more options, and this doesn’t even consider rivals to Kafka (Red Panda, Pulsar) We understand that this can feel overwhelming, so the sensible approach is to take a funnel approach and briefly compare lot of options to quickly to exclude and then narrow down to a few options for a deeper comparison.

Three of the biggest reasons we’ve come across for product exclusion are:

  • Too expensive for the budget
  • Bad fit for the team that will build and maintain the software (say, for example, knowledge of needing to use Java, in a team that only knows Python and SQL).
  • Doesn’t meet security requirements

So start here before comparing more technical requirements that will require more work to compare and make a list of must-haves in MOSCOW) before comparing nice-to-haves requirements (must haves in MOSCOW) before comparing nice-to-haves must have’s in MOSCOW) before comparing nice-to-haves to save on extensive comparisons.

Where to Start

Starting a product comparison can be the hardest part if you have no experience in the domain. There are a good few places to start: hiring expertise if you don’t have the time. Oakland is often called in to assess a client’s options. As we are tech-agnostic and have experience working with many different clients, we know what works and what doesn’t! (shameless plug for us), read blogs and books by subject matter experts (check out the Data Platform Journal another shameless plug) or use social media for advice.

Try to avoid straight-up sales pitches and salespeople early on; you are trying at this stage to gather information to make a better-informed decision later on, not make a decision right now. Most have useful guides and content which you can use to inform your decision-making.

Modern Data Stack is also a great website for finding a range of data solutions, though skewed towards smaller start-ups. Gartner is another option if you have a license, but it skews towards established buy solutions.

What to Watch Out For When Buying Software

Avoid bias towards slick sales staff and fancy user interfaces: vendors can hype up their products with years of sales experience, whereas engineers have little to no sales experience to hype up their build option.

Also, how much does a good User Interface (UI) for your users matter? For engineers, often not much, non-technical users, a lot.

Buyers may not realise that even heavily managed solutions require some configuration and support, not as much support as a build option, but this aspect is often forgotten about.

People tend to also forget about the 10% trap, where on average a complex solution requires 10% to 20% of its requirements to be met by highly customisable software.

We’ve seen a number of times people look at and/or test the simple use cases and forget the more difficult requirements they have, which inevitably means that the solution comes unstuck and can’t deliver on all of the requirements. Test your potential solutions against your most challenging must-have requirements.

Beware of your own Engineer’s Biases Towards Build

So far, this may sound a little biased towards “build“ solutions, so we’ll rein it in here.

Most experienced engineers would back themselves to build solutions that can also be bought if given the time and money, but just because you can doesn’t mean you should. Often, projects take longer and throw up difficulties along the way.

So, if the build and buy solutions cost roughly the same and both meet requirements, we would er on the side of caution and go for the buy option.

Another classic trap engineers fall into is pitching a solution that is technically better but doesn’t improve the organisation: making a pipeline run faster with fancy new tech doesn’t always make the organisation’s profits larger.

If you’re not an engineer, then you may be more likely to be biased towards a “buy“ option in our experience, though again, there are exceptions.

Summary

So, in summary: avoid extremes, gather information before engaging sales teams, compare using the most difficult and important requirements first, be creative and check your biases towards either buying or building.

Author: Jake Watson

The post Should you build or buy your data platform? appeared first on Oakland.

]]>
https://weareoakland.com/blog/should-you-build-or-buy-your-data-platform/feed/ 0
What is the environmental impact of your data? https://weareoakland.com/blog/what-is-the-enviromental-impact-of-your-data7624/ https://weareoakland.com/blog/what-is-the-enviromental-impact-of-your-data7624/#respond Mon, 18 Sep 2023 15:14:25 +0000 https://www.theoaklandgroup.co.uk/?p=7624 There is Something Human about Waste Let’s start by setting the scene – an all too familiar one. Humans love waste. We throw away around 100 billion pieces of plastic every year in the UK, and around 1.9 billion tonnes of food that can be eaten is discarded on an annual basis, and even on...

The post What is the environmental impact of your data? appeared first on Oakland.

]]>

There is Something Human about Waste

Let’s start by setting the scene – an all too familiar one. Humans love waste. We throw away around 100 billion pieces of plastic every year in the UK, and around 1.9 billion tonnes of food that can be eaten is discarded on an annual basis, and even on a household level – between 9-16% of energy is wasted on standby.

This fascinating graphic from No Planet B by Mike Berners-Lee shows the world supply chain for food based on energy loss at each stage. As you can see, at every level of the picture, our food waste. The depressing part can be seen in the bottom three rungs; even if we’re lucky enough to have food we still waste around 50% of it.

The impact of this level of overconsumption is evidenced in the world around us: fires in Maui, droughts in Uruguay, and extreme flooding in Bangladesh – all directly attributed to climate change driven largely by Greenhouse gas emissions.

But what does this have to do with data?

It doesn’t come as much of a surprise that physical human behaviour also maps onto digital human behaviour. The attitude towards waste and overconsumption has transcended from the physical to the digital in the case of data – and that has led to some huge inefficiencies in the way that we currently manage data – directly resulting in some pretty stark impacts on the environment. We talk about inefficiency through a few lenses – the first being the underutilisation of resources which we can see represented by the fact that the utilisation rate of on-premises servers is only around 18%, and it doesn’t get much better when we talk about cloud storage and usage, where around 33% of paid-for services are considered waste.

Combine this with the fact that 55% of stored data is considered dark (currently unused, with no future use-case) – a frightening statistic! That means stored and processed data that will never realistically be used for operational or analytical purposes. It’s one of the darker legacies of the big data wave of 2017!

Finally let’s not forget the onslaught of Artificial Intelligence. The amount of data we are creating is out of control.

Data Centres – The Lords of Darkness

When you consider the artificially inflated needs of server racks, the ramifications of this wastage are very real – data centres are being constructed at a rate of knots and are having pretty dire consequences on the environment.

Think about all the utilities that you need to maintain a data centre, primarily you might just think about the energy costs of powering your servers. Ignoring the energy requirements for cooling and lighting, server energy usage on its own has the potential for improvement. Server configuration and distribution of workloads can often drive higher than necessary energy usage.

Due to the sheer amount of energy digital technologies power through, data centres have become the world’s second largest source of greenhouse gases (2.7%) – behind petrochemicals but ahead of the poster boy of climate change, the airline industry (2%).

With no signs of slowing down, the environmental costs of increasing energy consumption will be around 14% of the world’s share by 2030.

All of this has led to data and digital becoming a primary concern for those interested in sustainability initiatives. Most large businesses require decarbonisation initiatives to meet regulatory requirements (e.g. SECR or CSRD). As the issue becomes more and more severe, these regulations are expanding from not only being focused on direct emissions (e.g. owned vehicles, building emissions) but also including scope 3 or supply chain emissions. In the context of data, on-premises data centres can utilise as much as 25% of an organisation’s total energy expenditure, so when you add your cloud usage into the picture – you can see the scope for improvement.

So what?

So we’ve set the scene – but why should businesses care? The way Oakland views this, there are four main drivers of why you should care as a business – boiled down into direct and indirect impacts. If you are considering or are currently involved in any kind of digital transformation, these are the metrics you should consider.

Cost

Direct impacts are associated with cost and regulation. With wastage comes excess fees – you can see enormous waste in resource use across cloud estates, and there is an associated cost with this resource wastage. Waste is the key word here – it’s cost that isn’t being spent on business outcomes.

The main concern with many sustainability initiatives is that they cost a lot of money and negatively impact business processes. However, with data, many of the issues are down to businesses not operating as efficiently as possible, meaning both environmental and financial costs are spiraling for everyone! By tackling these problems, you could potentially have an environmental sustainability initiative that maintains business processes and operations and positively impacts your bottom line.

Regulatory

Regulations are coming in that expect businesses (corporate sustainability reporting directive) to report not only their emissions for scope 2 and 3 (direct emissions and supply chain emissions) but also carbon reduction initiatives. You become more compliant if you can show how you plan to reduce carbon emissions across scopes.

Reputation

Reputation is critical – businesses are now becoming more attractive to prospective employees and customers because of their sustainability credentials. Imagine a data engineer who wants to go and work for two competing organisations. Both roles are bog-standard data engineering, but one of the organisations is committed to its environmental footprint by making its data infrastructure as carbon neutral as possible. The same goes for customers; if you can offer a sustainable option against one that hasn’t even thought about it, there’s a clear winner.

Doing what is right for the planet

Arguably THE most important driver – the long-term ecological disasters that are likely to increase due to the ever-increasing number of data centres and uncontrollable data estates. We mentioned them earlier, but on a more local level, we are seeing huge impacts of data centre development in places like Ireland (18% of energy is used by data centres); and in the US, where new data centres are being constructed in sunny, arid areas to take advantage of PV, we’re seeing reliant water consumption leading to severe droughts.

So really we need to tackle this at the level that we do business and change how we do things. Head to our Water Utilities page for more information on how we work with the water sector.

How to Approach Change

We need to talk about a new way to look at data operations. Can you support the delivery of both cloud and on-premises data estates in the most efficient possible way and have a more muted impact on the environment whilst having a more positive impact on your finances?

Let’s talk about FinOps

There’s been an obsession with FinOps in the past five years – and what that tends to focus on is cloud efficiency, but we saw from the stats before that it’s not necessarily working.

Rates of adoption are pretty low, action is limited to just IT or data teams, and cost becomes once again “a data problem.” When things are just “a data problem,” we don’t see the action which is needed.

So, let’s widen the conversation – How is sustainability baked into your data strategy? – that’s one of the key drivers to create buy-in to data across the business. As we’ve seen the next levels of both regulation and CSR being rolled into key strategic goals for organisations, we can leverage our approach to data in a way that appeals to these initiatives. Cut the waste and your cost can drop whilst meeting your strategic sustainability goals.

The answer? Include Sustainability

So, we approach data with a fourth lens – a lens of sustainability.

This acknowledges there is an issue with how we do business: by handling processes more efficiently we can save energy and carbon emissions, and by driving down energy spend and resource usage, cost savings can be achieved!

By revolutionising how we look at and assess the efficiency of our data estate, as a vehicle to support organisational sustainability initiatives we can drive the value from our data and drive the business outcomes needed from the data itself! This revolution moves us from looking at FinOps, to GreenOps.

Alas this isn’t easy – although we can quite easily say “how much” we spend financially on data, it is difficult to get a full picture of what our energy actually costs.

Oakland and our partners at Interact have devised a four-step process that businesses and data teams can use to start to adopt a GreenOps approach.

There are four main stages: Display, Diagnose, Decide, Deliver.

Display

You can’t take action without first understanding where you currently are – think of this as a baselining stage, but there are some major challenges in this phase – starting with understanding the impact your cloud is having. You need to consider many things, such as replication factors and networking – things that aren’t necessarily covered by existing cloud reporting.

Diagnose

A much more straightforward process – once you have completed your assessment, the next stage is to identify the key areas that are having the biggest impact without reaping benefits. One of the critical things from a cloud perspective could be where you do the majority of your computing, on-premises you want to look at utilisation rates and potential for consolidation to provide short term wins.

An example of this would be when we talk about the statistics of average utilisation of servers; combined with the linear nature of power usage vs. utilisation, you can triple the existing “utilisation” of that server (making it 54% utilised), whilst only using around 50% more energy.

That could result in a 300% increase in productivity for just a 50% increase in energy usage. Think about the carbon impact of this change and the cost savings that can be attached to your energy bills! With the high cost of energy, this is well worth an investigation.

Decide

Take the insights and diagnosis which enable you to start making decisions, reviewing what processes are necessary and if it is possible to change the location of your servers and the times of your heaviest usage.

Consider the wider impact that this has on your business – are there applications or data that you store on-premises that would be better served by migrating to the cloud? Or alternatively, you may even find potential for saving energy and money by moving your “archival” data back to on-premises servers.

The decisions you make here can be broken down into two lenses – short-term and long-term. Short-term decisions are those largely around consolidation and immediate fixes, such as changing the time of an ETL process to when a grid is being run on more renewable energy. To make these kinds of impacts, you don’t always have to fundamentally change what you’re doing in the short term – in fact, with Oakland’s carbon efficiency tool, we have seen that by shifting your heavy compute to one day rather than another has the potential to drop associated carbon emissions by around 10% – which costs nothing!

Your longer-term view of things might be migrating on-premises data to the cloud, but to do this most efficiently you need to understand the potential environmental impact – but the most important thing is to create a plan.

You want to focus on the areas you see the most potential to support your business and data strategies. If you’re going for a big push on governance, perhaps you need to initially consider where all that pesky dark data is sat and create a plan to decommission it, if you’ve got some horribly inefficient applications that aren’t optimised for the cloud – then create a plan to refactor them.

Deliver

Migrate, consolidate, and even delete! You need to deliver to the plan. When it comes to the delivery of these larger scale migration projects, you need to consider business continuity – how are you going to maintain services whilst doing this migration? One way to approach this is through introducing new tooling – for example, our partners at Starburst provide a distributed analytics engine that enables you to continually query from both on-premises and cloud instances to maintain services during a migration process.

What challenges might you face?

A lack of transparency from cloud providers on the carbon impact their services are making (for example, when we talk about replication factors, you need to consider that AWS lambda instances need to be replicated six times across different servers – meaning your carbon impact is six-fold).

A lack of publicly available data to support on-premises data assessments. Without support and data, it will take a long time to understand the existing impact that your large on-prem data centres are having.

We face a number of challenges as an industry, and to take on a GreenOps approach we need to be equipped with the tools and capability to assess and benchmark where we are, from an on-premises perspective and a consolidated view across both cloud and on-premises.

What are the next steps?

To support businesses in taking their next steps towards GreenOps, Oakland and Interact have partnered to create a revolutionary service based on the four-step method we’ve talked about here. Interact are experts in on-premises data centre efficacy assessments and how you can save money and energy by consolidating and reconfiguring servers.

Oakland can support organisations in reporting on, understanding, and investigating their cloud estate using our carbon efficiency tool – an enhancement from the well-regarded Thoughtworks cloud carbon footprint tool to create an estimate of the carbon footprint of your data estate.

We work together to identify how you can streamline your data estate and optimise for both carbon and cost perspectives. It’s important to note that the platform itself is only a part of the battle; you need governance to ensure that retention policies are being fulfilled, which can have a huge impact on your dark data, which needs to be baked into your data strategy as a means of alignment, and from an analytics perspective, you have to start to report on data sustainability as a KPI to support carbon reduction initiatives.

As an industry, we need to do better for ourselves and the planet by working together to minimise the environmental impact of our data in a cost-efficient way.

If you’d like more information about Oakland’s revolutionary GreenOps service please get in touch by emailing Luke.sharma@theoaklandgroup.co.uk

The post What is the environmental impact of your data? appeared first on Oakland.

]]>
https://weareoakland.com/blog/what-is-the-enviromental-impact-of-your-data7624/feed/ 0
Why Invest in Data Quality? https://weareoakland.com/blog/why-invest-in-data-quality/ https://weareoakland.com/blog/why-invest-in-data-quality/#respond Wed, 23 Aug 2023 09:31:22 +0000 https://www.theoaklandgroup.co.uk/?p=7551 This can seem like a rhetorical question: you should always invest in Data Quality! But we are still not investing enough: surveys show Data Quality issues are increasing in most organisations and on average, take up 34% of a Data Engineers time instead of them creating value by adding new features. This increases to 50% in large...

The post Why Invest in Data Quality? appeared first on Oakland.

]]>
This can seem like a rhetorical question: you should always invest in Data Quality! But we are still not investing enough: surveys show Data Quality issues are increasing in most organisations and on average, take up 34% of a Data Engineers time instead of them creating value by adding new features. This increases to 50% in large Data Platforms.

All these Data Quality issues add up, with bad Data Quality costing organisations on average $15mil a year.

Data Quality investment is also an investment in high-quality AI and ML, as you’ll likely get more accurate AI and ML results from improving Data Quality than changing your AI model and code.

Having Data Quality checks in place helps reduce “data downtime” for outages and fixes, which subsequently increases the overall reliability of the Data Platform. Highly reliable data leads to more trust in data and better-informed decision-making.

Better decision-making should increase profitability, productivity, and confidence of the whole organisation, which in turn usually leads to more investment in data and, as a result, going back to the start: Increasing the quality of data again.

All this creates a “Virtuous Cycle” of Data Quality, constantly improving your organisation:

If Data Quality decreases the opposite happens with a negative cycle.

Better Data Quality testing should also reduce the blast radius of the issue to a few Data Engineers rather than hundreds or thousands of users as more issues are being found earlier:

The fewer users impacted, the smaller the cost caused by the issue, which should again pay back any investment in Data Quality in large multiples.

Do I need a Data Quality Framework?

I know that setting out to build a framework for anything requires time, and you’ll have many competing concerns, so we understand if you feel reluctant to build one, especially if you are a small team with a limited budget.

But there comes a point where fighting lots of local battles with Data Quality becomes more inefficient than building out a framework to reduce Data Quality issues over the long term.

We’re not going to do a deep dive on Data Quality frameworks here, as they are often tied to wider Data Governance frameworks (you can download our guide here) We will say that whatever framework you use, make sure it’s cyclical so that it’s always improving and you are acting on any emerging issues in a timely manner.

How Do I Test for Data Quality?

Classically, Data Quality tests are a set of rules that test between the actual and desired state of data. The desired state may not be perfect, but ‘good enough’. What counts as ‘good enough‘ varies from dataset to dataset, which makes Data Quality more challenging.

What do we normally test in a dataset, though? The DAMA International’s Guide to the Data Management Body of Knowledge says there are six dimensions to Data Quality:

  • Accuracy: does the date look how we expect it to?
  • Completeness: are there any unexpected missing values?
  • Uniqueness: no duplicates!
  • Consistency: does a person’s data match in two different datasets?
  • Timeliness: is the data out of date?
  • Validity: does the data conform to an expected format? Think postcodes, emails, etc.

Tracking all these dimensions for every dataset at every stage of your pipeline is a lot of work, probably too much work. Therefore, a trade-off is often required to focus on areas where Data Quality will have the most impact on the business.

You also have to beware of false positives or minor issues being blown out of proportion, overwhelming your engineers with too many issues. It can help if you categorise your Data Quality issues by severity just like other software issues.

You also have to take into account the mental wellbeing aspect too: few Data Engineers and Analysts want to spend a large percentage of their time fixing Data Quality issues over a long period of time.

There is help, though, with software frameworks to help you write Data Quality testing:

Most of the above profile your data and setup recommended tests for you to use, saving you some time configuring them yourself.

But in reality, we see a lot of custom-made Data Quality testing, partly because Data Quality struggles for investment, so it is usually done in an organic, ad-hoc manner.

The above products work best in development and staging environments, so you can find issues before they enter production or use them as circuit breakers, to stop a Data Pipeline if the incoming or outgoing data is of poor quality.

It is also worth mentioning that you can use constraints in a Warehouse or Lakehouse schema, which have the benefits of not requiring another software library but are not as feature rich (you will likely have to setup your own notifications for alerting).

It is also important to inform your users of any Data Quality issues as soon as possible so they don’t waste time finding out for themselves or use data that is untrustworthy. This can be done through notifications and alerts, though we’ve also had a lot of success creating Data Quality dashboards that sit alongside existing reports and can be easily referred to by users.

Latest Concepts in Data Quality

There has been significant innovation in Data Quality in the last few years, so we present below the concepts to take your Data Quality process to the next level.

This will require more investment, but it will give you an edge over your competitors to make better informed decisions as you’ll have more trustworthy data. This investment should also pay back long term with less time wasted fixing Data Quality issues.

What is Data Reliability and do I Need it?

Data Reliability gives Data Quality more of a support focus, which makes sense as most Data Quality issues in production will be dealt with as a support issue to a Data Platform.

Data Reliability takes a lot of its thinking from Site Reliability Engineering (SRE), which treats support as more of a engineering problem, where you examine your past and current support tickets and look to decrease them with engineering or better processes.

With Data Reliability, you would look to get a baseline of Data Quality issues per week or month and then look at ways to reduce them and monitor to see if the changes have reduced the number of issues and/or reduced the amount of time spent on issues.

The changes to improve Data Reliability can be technology-based:

  • New or updated tooling
  • Better automation of when a pipeline fails or automated actions to respond to a data issue

Or the changes can be process-oriented:

  • Writing better documentation to avoid common issues
  • Incident playbooks so the whole team can more quickly respond to a issue in an consistent way.

You can rather cynically say Data Reliability is just Data Quality with a feedback loop and a time series graph, but it is there to make sure you avoid short term thinking about Data Quality and instead consider long term improvements that will make your data platform more efficient and trustworthy.

Data Reliability Cycle

You may also set targets such as “99.9% of data will refresh on time” or “A maximum of 33% of engineer time should be spent on support issues“ as well. As mentioned before, it can be impossible to achieve perfect Data Quality, so aiming for a reasonable target instead can avoid engineer burnout.

What is Data Observability and do I Need it?

Data Observability is about gaining a Data Platform or organisation-wide understanding of your Data Quality.

It arguably goes beyond Data Quality by adding metadata features normally found in a Data Catalog: cataloguing schemas of datasets and data lineage. These features allow you to more quickly find a Data Quality issue by tracing the lineage of the issue and also you gain the ability to see how much Data Quality is impacting your organisation.

Data Observability software can often also come with Machine Learning (ML) algorithms to detect anomalies in data, so you can be warned about issues you haven’t even thought of yet.

We’ve seen products either extend a Data Quality framework with Data Catalog features such as Monte Carlo and Big Eye. Or existing Data Catalogs add Data Quality functionality, such as Datahub, which imports Data Quality tests created by Great Expectations and dbt tests. Both Soda and Monte Carlo have integration with the Data Catalog Alation.

What are Data Contracts and do I Need Them?

Data Contracts make a contract between a data producer and a data consumer, so the consumer knows what data to expect from the producer.

While you can replicate some of a Data Contract’s benefits by tracking the schema of the data produced, a Data Contract is meant to go beyond that by giving you a full suite of metadata about the data:

  • The data’s schema.
  • How the data is calculated.
  • Who owns the data?
  • What is the data lineage?
  • How to access the data.
  • What is the data’s expected quality, availability, etc.
  • Plus anything else that is relevant to the data.

You may think Data Contracts are redundant if you have a well-maintained Data Catalog, as they capture similar information, but Data Contracts are designed to be checked during every run of a Data Pipeline and have some action in the pipeline if the Data Contract is broken:

  • Stop the pipeline with a circuit breaker.
  • Alerting.
  • Moving data that doesn’t meet the contract to a manual checking table.

For an example, Paypal has open-sourced their Data Contract template.

Data contract schema

https://github.com/paypal/data-contract-template

This should create more positive collaboration between data producers and consumers because they have a collective agreement of what the data should look like. It is not uncommon to have a poor working relationship where a producer makes changes without telling consumers or consumers accessing data in way not recommended by the producer.

One issue with Data Contracts is that they are a new concept, so require more work to implement at present, though that will likely change in the near future as more companies adopt them.

Most of the examples of Data Contracts we’ve seen so far use Apache Flink and the Kafka Schema Registry, so assume you are using streaming, though there are some examples that use batch processing.

Data Governance and Data Quality

Good Data Governance can also improve quality of data. It is important to know where data is coming from, who owns it, for what purpose data is being transformed, and finally, what is the impact of poor availability and data quality: all helped by having Data Governance properly implemented.

Some of the above concepts (Data Contracts and Data Observability) can also improve Data Governance, so investing in Data Quality can also be an investment in good Governance too.

How Does This All Fit Together?

The diagram below is one example of how it all fits together:

  • Any code changes are tested in development and/or test environments with Data Quality Tests to check that any changes won’t have a negative impact on Data Quality.
  • Source Data at the start of the data pipeline is checked to see if the Data Contract is held; if not, a circuit breaker may kick in, stopping the data pipeline early to avoid processing unsuitable data.
  • Data Quality tests are also run in production, which can feel like duplication from testing in development, but there may be changes caused by moving to a production environment (different data, etc.).
  • Data is collected for observability checks by Data Observability software, looking for any anomalous data: a department budget that goes from £10k to £1mil or 10x increase in rows for a table, for example. This can replace a lot of tests, but not all of them.

You’ll also be collecting Data Quality metadata to improve your Data Reliability.

Making all this work together seamlessly isn’t cheap and will take time, but as mentioned, poor Data Quality will also cost an organisation a lot of money. So we recommend tackling this in an agile manner by improving Data Quality in small increments, one change at a time, starting where it will have the most impact.

Summary

Data Quality is a difficult subject to tackle, due to it being a slightly different problem in every organisation and never “perfect”. That said, there are lots of options to help improve the quality of your data, so you should be able to get to “good enough“ if you give Data Quality enough priority and forethought.

Jake Watson is a Principal Engineer at Oakland

Get In Touch 


The post Why Invest in Data Quality? appeared first on Oakland.

]]>
https://weareoakland.com/blog/why-invest-in-data-quality/feed/ 0