Tech Talk | Oakland Tue, 16 Dec 2025 10:55:17 +0000 en-GB hourly 1 https://wordpress.org/?v=6.9.4 https://weareoakland.com/wp-content/uploads/2024/01/cropped-oakland-favicon-150x150.jpg Tech Talk | Oakland 32 32 Microsoft Ignite 2025: The Debrief from Oakland’s LinkedIn Live https://weareoakland.com/blog/ms-ignite-2025-debrief/ Tue, 16 Dec 2025 10:54:06 +0000 https://weareoakland.com/?p=9853 If you missed our very first LinkedIn Live, don’t worry - we’ve pulled out all the key points in this handy blog. Here are the practical takeaways from our four experts as they looked back at Microsoft Ignite 2025 and what it all means for the future of data and AI for your organisation!

The post Microsoft Ignite 2025: The Debrief from Oakland’s LinkedIn Live appeared first on Oakland.

]]>
If you missed our very first LinkedIn Live, don’t worry – we’ve pulled out all the key points in this handy blog. Here are the practical takeaways from our four experts as they looked back at Microsoft Ignite 2025 and what it all means for the future of data and AI for your organisation!

The panel included:

  • Andy Crossley, Oakland’s Chief Technology Officer (and our man on the ground at Microsoft Ignite 2025)
  • Rajan Chavda, Oakland’s Microsoft Fabric expert
  • Chris Hill, Oakland’s Purview and Microsoft Partnership Lead
  • MLG (Mike), Oakland’s Principal AI Engineer

In typical Oakland style, nothing in here is theory or hype. It’s a grounded reflection on what Microsoft announced, what actually matters for organisations, and how these changes affect real data teams today.

Prefer to watch the full session? No problem – view the recording below!

What it Felt Like to Be at Microsoft Ignite 2025

Andy opened the session by sharing his experience of being on the ground in San Francisco. The scale was huge and the pace frenetic. Instead of a single moment where Microsoft revealed something dramatic, it became clear that innovation now moves so fast that Ignite is more of a checkpoint than a product launch.

What Andy heard most people talking about was the concept of the Frontier Firm. It’s Microsoft’s attempt to describe the kind of organisation that thrives in an AI driven world. Whether people loved the term or not (which is the subject of some debate), it dominated the conversations across the conference. The idea is not that you become a futuristic organisation overnight. Instead, every organisation will move at its own pace, building its own version of what Microsoft describes.

Understanding the Frontier Firm

What Microsoft Means

Microsoft’s model proposes three stages:

  1. Humans supported by AI and Copilots
  2. Humans and agents working together
  3. Humans overseeing many agents that run key operations

The focus is not on shiny robots or science fiction. It’s on how work gets done and how AI can start taking on meaningful chunks of operational activity so people can focus on higher value work.

What this Means in Practice

The Governance Challenge

Chris highlighted that Microsoft is predicting more than one billion three hundred million agents by 2028. That scale raises the same problems we have always had in data. If you do not control access, permissions, and visibility, you quickly end up with an uncontrollable mess.

The Everyday Work Challenge

Rajan looked at the practical side. For him, the most important question is how AI reduces repetitive or low value tasks. The point is not to create agents for the sake of it. It is to support the real day to day work that teams do.

The Technical Reality

MLG spoke from his experience building agentic systems. Agents are powerful, but they can also be unpredictable. What stood out for him at Ignite was the improved tooling across Microsoft Foundry and the new Agent 365 ecosystem. These offer developers better ways to build, test and monitor agents so they can behave consistently inside real organisations.

The Cultural Reality

Andy added that organisations need confidence and clarity. Most people understand simple Copilot features. Most can imagine advanced automation. The challenge is everything in between. Organisations need a roadmap that is ambitious but also realistic.

Adoption Barriers and Opportunities

A question during the live event asked about the biggest challenges to adoption.

The group agreed that organisations need all three of the following:

  1. Technical Readiness

You must have a clean, well governed and discoverable data estate.

  1. Cultural Readiness

People must understand when and how to use agents and what good looks like.

  1. Governance Readiness

You cannot deploy agents at scale without policies, controls and monitoring.

As Andy said, you would not hire a thousand people without onboarding and training. The same applies to AI agents.

MS Ignite Announcements by Product

The team reviewed the major changes across Fabric, Foundry, Purview and Microsoft AI services.

Microsoft Fabric: The Intelligence Layer Takes Shape

Rajan took the audience through the biggest changes in Fabric.

Fabric IQ

Fabric IQ introduces a new layer of intelligence across the platform. Ontology maps let you define the real business relationships behind your data. This makes insights richer and gives agents more meaningful context to work with.

Operations Agents

For Rajan, Operations Agents are the standout feature. They allow organisations to close the loop between insight and action. Instead of dashboards hoping someone does something, Fabric can now monitor data, suggest actions, and trigger workflows. This is surfaced directly in tools like Teams.

Interoperability with Databricks and Others

One of the most encouraging themes was Microsoft’s shift toward open integration. Fabric will work more closely with Databricks, Snowflake, and SAP. OneLake-backed compute for Databricks is planned for future releases. Andy noticed how often interoperability came up at Ignite. Microsoft now accepts that real organisations run mixed estates.

Purview and Agent 365: Governance Grows Up

Chris explained why Purview had a smaller public presence at Ignite. The reason is that governance is being repositioned within the new Agent 365 framework.

Agent 365 Brings Five Key Capabilities

1. A registry of every agent in your organisation

2. Central access control using Entra ID

3. Visibility into what agents do

4. An understanding of how agents interact across departments

5. Security anchored in Purview

This directly addresses the problem of agent sprawl. As Chris put it plainly, AI governance is not optional. It is essential.

Microsoft Foundry: A Mature Space for AI Developers

MLG walked through how Foundry has evolved into a unified space for building, evaluating and managing agentic systems.

Foundry IQ

This gives developers the ability to benchmark agents, measure accuracy and check factual grounding. Without testing, there is no trust. Foundry IQ is Microsoft’s answer.

Agent Lifecycle Management

Foundry now supports versioning, AB testing and controlled retirement. If one agent outperforms another, you can replace it safely and consistently. This is vital if organisations want agents to work at scale.

What Does All this Mean for Organisations?

From the discussion, four clear messages emerged:

  1. Strategy First

Technology should support your business strategy. Without clarity on what you are trying to achieve, AI becomes a distraction.

  1. Readiness Matters

Adoption will not work if your data estate, governance or culture is not prepared. Just because the capabilities exist does not mean you are ready to use them.

  1. Trust Must Be Earned

People will only trust agent driven actions if results are reliable, transparent and grounded in fact.

  1. Governance will Decide Success

Agent 365 and Purview will be central to safe adoption. Governance gives you the confidence to scale without losing control.

Final Thoughts from the Oakland Team

Ignite 2025 didn’t feel like a tech expo where one big announcement stole the show. Instead, it marked a point where Microsoft began stitching together years of change into a more coherent story.

The frontier firm is not a single destination. It is a direction of travel. Organisations will move at different speeds and in different ways. What matters is strong data foundations, clear governance and a realistic roadmap.

Oakland’s focus remains the same. We help organisations build the maturity they need to use AI responsibly and effectively. With the right support, AI and agents can streamline operations, free teams from repetitive tasks and unlock new value in the business.

FAQs about Microsoft Ignite

What is Microsoft Ignite?

Microsoft Ignite (MS Ignite) is the annual event where Microsoft showcases the trends, innovations, and future vision for their productivity, cloud, and security technologies. The event takes place over several days and is packed with keynotes from 2,000+ speakers, 800+ interactive sessions, networking opportunities, and more. It’s attended by over 20,000 people in-person, with a further 200,000 joining digitally.

Who should attend MS Ignite?

Developers, IT and data engineers, cloud architects, business leaders (including startup founders) will all find the event informative and inspiring. In the words of Microsoft themselves: ‘Microsoft Ignite 2025 isn’t “just another conference.” It’s a unique gathering that brings together technology leaders, tech professionals, developers, founders, and Microsoft partners for four days of immersive learning and groundbreaking announcements.’

When is Microsoft Ignite 2026?

The next event will be held at the Moscone Center in San Francisco between 17th and 20th November 2026.

The post Microsoft Ignite 2025: The Debrief from Oakland’s LinkedIn Live appeared first on Oakland.

]]>
How Microsoft Fabric is Reshaping the Enterprise Data Platform https://weareoakland.com/blog/microsoft-fabric-enterprise-data-platform/ Fri, 30 May 2025 12:40:52 +0000 https://weareoakland.com/?p=9546 At a recent Oakland event co-hosted with Microsoft, we had the pleasure of welcoming Chris Webb, a seasoned Microsoft Fabric expert and member of the Fabric Customer Advisory Team. Chris ran a session exploring Microsoft Fabric in full, answering: Below, we round up the insight Chris shared during the informative session – a must-read if...

The post How Microsoft Fabric is Reshaping the Enterprise Data Platform appeared first on Oakland.

]]>
At a recent Oakland event co-hosted with Microsoft, we had the pleasure of welcoming Chris Webb, a seasoned Microsoft Fabric expert and member of the Fabric Customer Advisory Team. Chris ran a session exploring Microsoft Fabric in full, answering:

  • What is Microsoft Fabric?
  • Why does it exist?
  • Microsoft Fabric VS Power BI: Is there a need for both?
  • Microsoft Fabric VS Databricks: How do they stack up?
  • Why does Fabric promise to reshape the landscape of enterprise data platforms?

Below, we round up the insight Chris shared during the informative session – a must-read if you’re considering modernising your data stack. 

How Microsoft Fabric Empowers Radical Simplification

For organisations navigating increasingly complex data estates, Microsoft Fabric promises a radical simplification. 

“Fabric isn’t just a bundle of tools; it’s an integrated platform designed from the ground up to work seamlessly, making it far easier to extract value from data without the usual configuration headaches.”

– Chris Webb, Microsoft Fabric expert and Fabric Customer Advisory Team Member

But what spurred Microsoft to develop Fabric, and what problems did they set out to solve with another data platform?

The Need for Another Microsoft Data Platform

If you’re already invested in a secure Azure data platform, it may not be the first time you’ve asked this. Azure already boasts tools like Power BI, Azure Data Factory, Spark, and Synapse. However, as Chris noted, combining these services to build a functioning data platform has historically been complex and time-consuming. 

Interested in Azure? Find out how we delivered a Microsoft Azure platform to leverage market expansion and growth for Emerald Publishing.

What Fabric offers is a unified software-as-a-service (SaaS) platform where everything is pre-integrated. With Fabric, you don’t need to spend time wiring services together, managing multiple security layers, or stitching storage solutions across products. You simply turn on Fabric, and everything, from data ingestion and transformation to data warehousing, reporting, and governance, works together by default. 

At the heart of Fabric is OneLake, a single storage location where all workloads store data in the Delta format. This allows teams to move from ingestion to insight without needing to copy or reformat data, a significant leap forward in usability and performance. 

Read our blog to understand if it’s right for you:

Microsoft Fabric: Power BI with Superpowers

For many, Fabric’s appeal begins with familiarity. Chris, a Power BI specialist at heart, explained that Fabric ‘is effectively Power BI with superpowers’. It builds on Power BI’s ease of use and widespread adoption, empowering 30 million monthly users and over 375,000 paying customers with enterprise-grade tools for engineering, science, and real-time intelligence. 

In part, Microsoft Fabric was built because so many organisations already trust Power BI. Instead of talking about Microsoft Fabric VS Power BI, we’re talking about the two solutions learning from one another.

“We’re taking the Power BI way of working (focused on customer feedback and monthly innovation) and bringing it to the enterprise data platform as a whole.” 

– Chris Webb, Microsoft Fabric expert and Fabric Customer Advisory Team Member

Oakland’s Partnership with Microsoft

Did you know Oakland is a Microsoft Data and AI solutions partner? Our people are Microsoft-certified in Data and AI, Power Platform, Infrastructure and Security, as part of the Microsoft Partner Network.

Rapid Growth and Real-World Adoption of Fabric

Since its launch just 15 months ago, Fabric has seen explosive adoption. Over 19,000 customers are already using Fabric, and more than half of them are running three or more workloads beyond Power BI. What’s more, 70% of Fortune 500 companies are now Fabric users. 

Chris highlighted UK-based case studies to show real-world momentum:

  • Iceland, the supermarket chain, is using Fabric for real-time transaction analysis.
  • Centrica, a major energy company, and the London Stock Exchange Group are deploying Fabric to power their enterprise data strategies.

Who will be next?

What’s New from FabCon: Key Feature Announcements 

We couldn’t be in Las Vegas for FabCon, but Chris brought the pizazz to Leeds and touched on some of the most exciting product announcements and updates. 

OneLake Security advancements

Now, you can apply row-level and column-level security once in OneLake, and that single policy flows through all workloads. That’s whether users are querying data in Python notebooks, SQL, or viewing dashboards in Power BI. It’s a game-changer for data governance, significantly reducing complexity for enterprise teams managing sensitive data.

Copilot for Power BI sees accessibility expanded 

Previously limited to higher-capacity SKUs, Copilot for Power BI will now be available in all Fabric capacities starting from F2. This means organisations of any size can now use natural language to explore data, create reports, and uncover insights without writing a single line of code. 

Chris confirmed that by the end of this month (May 2025), the feature will be much more accessible, democratising analytics powered by artificial intelligence

Real-time intelligence gets a boost 

Real-time analytics has emerged as a breakout feature in Fabric. Chris described how organisations, like Porsche Racing, are ingesting high-frequency telemetry data from vehicles and analysing it in near real time. 

Recent updates to Event Streams and Event House now make it even easier to integrate and act on data from a range of sources, including new connectors, like MQTT and weather feeds. The ability to run SQL transformations in-stream was also highlighted as a major feature coming soon.

Materialised Views in Spark 

Chris demonstrated how Materialised Views allow Spark developers to build complex transformation pipelines: bronze to silver to gold layers without managing orchestration. Fabric:

  • Auto-resolves dependencies
  • Runs quality checks (read our article on why you should invest in data quality)
  • Surfaces visual lineage to simplify debugging

These views appear as Delta tables and can be queried directly in Power BI or SQL, showcasing Fabric’s tightly integrated architecture. 

End-to-end developer experience 

Another demo showed a deeply integrated development workflow combining SQL, Python, user-defined functions, and new variable libraries. These libraries make it easy to define parameters that update automatically when moving from dev to test to production environments, streamlining deployment and reducing the risk of errors. 

The notebook experience in Fabric now rivals the flexibility of platforms like Databricks, while offering much tighter integration with Microsoft tools. 

Use our blog to understand Microsoft Fabric VS Databricks: Taming your data assets with Databricks.

Will Power BI be Forgotten?

With all these advancements in mind, there was a common concern among event attendees: Has Power BI been left behind as Fabric takes centre stage? 

No.

Power BI continues to evolve as a core part of Fabric, with major updates to visuals, slicers, and storage modes. One of the biggest innovations is DirectLake mode, which combines the speed of import with the freshness of direct query. No more waiting for scheduled refreshes. Good news all-around!

Microsoft Fabric is More than a Product – it’s a Strategy

Chris Webb’s presentation made one thing obvious: Microsoft is positioning Fabric not just as another data tool, but a fundamental rethinking of how enterprises manage, analyse, and act on their data. I.E., their overarching data strategy.

From real-time intelligence to AI-driven development, unified governance to Power BI integration, Fabric empowers businesses to do more with less complexity. 

Head to our guide for full details on how to write your data strategy.

Modernise your data stack with Oakland

If your organisation is looking to modernise its data stack, now might be the time to consider Microsoft Fabric. Engaging a Microsoft Fabric consultant, like our experts at Oakland, can help you accelerate your adoption, avoid common pitfalls, and make sure your data investments deliver value from your data assets. 

Get in touch to talk about all-things data stacks and Fabric, or book a free exploratory workshop with our data consultants.

The post How Microsoft Fabric is Reshaping the Enterprise Data Platform appeared first on Oakland.

]]>
Should you build or buy your data platform? https://weareoakland.com/blog/should-you-build-or-buy-your-data-platform/ https://weareoakland.com/blog/should-you-build-or-buy-your-data-platform/#respond Wed, 08 Nov 2023 13:36:50 +0000 https://www.theoaklandgroup.co.uk/?p=7790 “Build Versus Buy” is More a Scale of Options than a Binary Decision As consultants, we’ve been involved in helping our clients to decide which is the best software to invest in (or not invest in some cases). This is quite a responsibility as we are putting our reputation in the hands of a vendor,...

The post Should you build or buy your data platform? appeared first on Oakland.

]]>
“Build Versus Buy” is More a Scale of Options than a Binary Decision

As consultants, we’ve been involved in helping our clients to decide which is the best software to invest in (or not invest in some cases). This is quite a responsibility as we are putting our reputation in the hands of a vendor, and we know that tech vendors are amazing at making claims. We see it in every trade show we attend. Therefore, it can be impossible to verify all the features of every product.

From our experience, this can be difficult to get right. So here are a few lessons in buying software (or not) to avoid making costly mistakes.

Limiting Your Potential Options is Sensible but Dangerous

One sales tactic is to present two options. This makes sense in some ways: management is time constrained and doesn’t want to assess all 100+ options that are possible, but this can lock you into a limited way of thinking, especially if you only assess two extreme “buy“ and “build” options.

Diagram to show investments in software for a data platform on a scale

Scale of Product Types from Build to Buy. By Jake Watson.

So you want to look at a variety of options, and for most organisations, look at options across the scale. We find a managed Platform as a Service (PaaS) in the middle of the scale, fits well for most but not all!

Think Beyond a Product

You may also want to consider a mix of both build and buy components; this can get you closer to fulfilling a set of complex requirements in a reasonable budget and timeframe.

One of our favourite examples is Fivetran, a vendor that offers a super quick and easy way of getting source data into your Lakehouse or Warehouse. However, it can cost a lot at scale, so you may end up doing the math and finding that building your own ingestion is more cost-effective for some, but not all, sources. Or try open source (AirbyteMeltano).

And yes, maintaining two or more systems is generally worse than one, but that should be a guideline, not a rule, as sometimes, you are the exception.

There are Too Many Choices Though!

You can go too far the other way and spend months going through dozens of options out of 1400+ products that exist in data and AI.

One example of the variety of options available to you is real time data processing, where you can choose from least to most managed:

*For those who haven’t come across IaaS, PaaS and SaaS we recommend this article.

We’ve probably missed a few more options, and this doesn’t even consider rivals to Kafka (Red Panda, Pulsar) We understand that this can feel overwhelming, so the sensible approach is to take a funnel approach and briefly compare lot of options to quickly to exclude and then narrow down to a few options for a deeper comparison.

Three of the biggest reasons we’ve come across for product exclusion are:

  • Too expensive for the budget
  • Bad fit for the team that will build and maintain the software (say, for example, knowledge of needing to use Java, in a team that only knows Python and SQL).
  • Doesn’t meet security requirements

So start here before comparing more technical requirements that will require more work to compare and make a list of must-haves in MOSCOW) before comparing nice-to-haves requirements (must haves in MOSCOW) before comparing nice-to-haves must have’s in MOSCOW) before comparing nice-to-haves to save on extensive comparisons.

Where to Start

Starting a product comparison can be the hardest part if you have no experience in the domain. There are a good few places to start: hiring expertise if you don’t have the time. Oakland is often called in to assess a client’s options. As we are tech-agnostic and have experience working with many different clients, we know what works and what doesn’t! (shameless plug for us), read blogs and books by subject matter experts (check out the Data Platform Journal another shameless plug) or use social media for advice.

Try to avoid straight-up sales pitches and salespeople early on; you are trying at this stage to gather information to make a better-informed decision later on, not make a decision right now. Most have useful guides and content which you can use to inform your decision-making.

Modern Data Stack is also a great website for finding a range of data solutions, though skewed towards smaller start-ups. Gartner is another option if you have a license, but it skews towards established buy solutions.

What to Watch Out For When Buying Software

Avoid bias towards slick sales staff and fancy user interfaces: vendors can hype up their products with years of sales experience, whereas engineers have little to no sales experience to hype up their build option.

Also, how much does a good User Interface (UI) for your users matter? For engineers, often not much, non-technical users, a lot.

Buyers may not realise that even heavily managed solutions require some configuration and support, not as much support as a build option, but this aspect is often forgotten about.

People tend to also forget about the 10% trap, where on average a complex solution requires 10% to 20% of its requirements to be met by highly customisable software.

We’ve seen a number of times people look at and/or test the simple use cases and forget the more difficult requirements they have, which inevitably means that the solution comes unstuck and can’t deliver on all of the requirements. Test your potential solutions against your most challenging must-have requirements.

Beware of your own Engineer’s Biases Towards Build

So far, this may sound a little biased towards “build“ solutions, so we’ll rein it in here.

Most experienced engineers would back themselves to build solutions that can also be bought if given the time and money, but just because you can doesn’t mean you should. Often, projects take longer and throw up difficulties along the way.

So, if the build and buy solutions cost roughly the same and both meet requirements, we would er on the side of caution and go for the buy option.

Another classic trap engineers fall into is pitching a solution that is technically better but doesn’t improve the organisation: making a pipeline run faster with fancy new tech doesn’t always make the organisation’s profits larger.

If you’re not an engineer, then you may be more likely to be biased towards a “buy“ option in our experience, though again, there are exceptions.

Summary

So, in summary: avoid extremes, gather information before engaging sales teams, compare using the most difficult and important requirements first, be creative and check your biases towards either buying or building.

Author: Jake Watson

The post Should you build or buy your data platform? appeared first on Oakland.

]]>
https://weareoakland.com/blog/should-you-build-or-buy-your-data-platform/feed/ 0
How to Harness Generative AI for Data-Driven Decision-Making https://weareoakland.com/blog/harnessing-generative-ai-for-data-driven-decision-making/ https://weareoakland.com/blog/harnessing-generative-ai-for-data-driven-decision-making/#respond Tue, 10 Oct 2023 14:17:15 +0000 https://www.theoaklandgroup.co.uk/?p=7701 “In the boardroom, one fact remains constant: data is the lynchpin of contemporary business. How many organisations can genuinely claim to have fully operationalised their data assets, though? If your data strategy still resides on the periphery of business operations, rather than being its driving force, you’re in good company.”  That illuminating passage was written...

The post How to Harness Generative AI for Data-Driven Decision-Making appeared first on Oakland.

]]>
“In the boardroom, one fact remains constant: data is the lynchpin of contemporary business. How many organisations can genuinely claim to have fully operationalised their data assets, though? If your data strategy still resides on the periphery of business operations, rather than being its driving force, you’re in good company.” 

That illuminating passage was written by ChatGPT (should we be surprised that even artificial intelligence (AI) calls for cleaner data?). Yet, while amusing, and with many of us having had a good play around with the technology, it really is time to consider generative AI as a key strategic asset for your business. 

Generative AI isn’t just another buzzword to add to your corporate vocabulary—it’s at the forefront of intelligent decision-making. You may already be familiar with analytical AI technologies like machine learning, but generative AI offers something more: the ability to create new, actionable insights by synthesising vast realms of data through custom AI systems.

What Are Some Key Gen AI Enterprise Use-Cases?

There are thousands of generative AI use cases out there – every organisation and its challenges are different after all. Yet, these are some of the most common that our AI consultants have worked with  over the past few years:

  • Automated Customer Interactions: Generative AI can power customer service solutions that offer unprecedented personalisation while streamlining operations. Recent surveys indicate that customers in some sectors prefer AI’s efficient and targeted responses over human responses.
  • Gen AI in marketing: GenAI suddenly makes life in the Marketing world much less resource and time-intensive. If you can think it, you can produce it! If you want a celebrity in your marketing campaign, no problem. For Nike’s 50th anniversary, we could see Serena Williams and a younger version of herself battling it out on the tennis court. Avatars can now also provide personalised and interactive messages. Imagine running campaigns that are not only data-driven but also continuously optimise themselves. Generative AI can produce innovative, personalised marketing content at scale.
  • AI in healthcare: Generative AI has dramatically reduced drug development and medical research cycles in the pharmaceutical and healthcare sectors. Accelerated healthcare research isn’t just about speed; it’s about enabling new avenues of research that were previously unthinkable.
  • AI for data governance: In a recent use case, a financial institution substantially leveraged Generative AI to improve its data quality. By creating simulated but realistic data sets, the institution could test the robustness of its fraud detection algorithms under various conditions before actual deployment. This enabled them to pinpoint weaknesses in their governance framework, tighten control mechanisms, and gain more reliable risk assessment insights.

What are the Challenges of Generative AI? 

In order for it to provide return on investment, generative AI brings with it several distinct challenges that businesses must be vigilant about: 

  • Lack of data integrity that produces untrue or biased outcomes
  • Unethical use that may lead to societal bias or human rights infringements
  • Copyright and intellectual property infringements and opportunities
  • Security and privacy rights
  • ESG Impact

Although there is currently no explicit law or legal framework to regulate AI use, the EU and UK are intent on releasing these soon. What’s more, there’s no getting around the fact that underlying data assets need to be fit for purpose to set foundations for compliant tools.  

How can AI be used in Data Governance?

In 2023, Gartner polled 2,500 executives, asking what the primary focus for GenAI initiatives in their businesses had been. Customer experience and retention (38% of all initiatives), and revenue growth (26%) came out on top, followed by cost optimisation (17%) and business continuity (7%).   

However, all this begs the follow-up question: how well did these initiatives do? 

A decade ago, dashboards were considered the pinnacle of data-driven decision-making. The limitations, however, became apparent when executives realised that the quality of underlying data was often inadequate. For AI implementations, the governance requirements are similar but exponentially more complex. Businesses must ensure that data governance policies are robust enough to manage the capabilities and risks of AI-driven decision-making processes.  

Bluntly put,  “Garbage in – Garbage out”. 

AI tools need to be managed as data tools. To make them fit for purpose, companies must have rigid oversight of the processes, policies, and governance of the data involved (both at ingress and egress) and the tool’s performance itself.  

An AI data governance framework should include: 

  • Data Quality and Integrity: Ensuring the data is unbiased and accurate.
  • Accountabilities and owners: For the data itself but also the development specifications and the outcomes.
  • Transparency: Where is the data being sourced from, how are the outcomes being used, and who can access these?
  • Ethical use: Does it have a negative impact, or is it a threat to people’s security and rights?
  • Human Oversight: AI doesn’t and shouldn’t fix its data problems to ringfence what the tool can and cannot produce, as hallucinations of the AI can become the norm.

Is your organisation ready for generative AI? 

There are four key questions every organisation should ask itself before embarking on a generative AI project:

  • Do you understand your business capability landscape adequately to scope where AI can improve your business performance?
  • Do you have a clear map of your information flow and data dependencies?
  • Does your business have a working Data Strategy or is it currently just a document sitting on the intranet?
  • Are your data platforms up to the task?

If you can answer yes to these questions, piloting Generative AI initiatives could be the next logical step. In today’s data-driven world, harnessing the power of AI is no longer an option; it’s a necessity. 

Generative AI is not just a buzzword; it’s a game-changer. It can transform how you handle data, making your operations smarter, faster, and more efficient. Whether you’re looking to automate tasks, enhance decision-making, or innovate your products and services, Generative AI is the key to unlocking these opportunities.

The next important question, though, is: where does Generative AI fit into your unique data landscape?

How Oakland can help you drive value from generative AI

Our experts understand how Generative AI can seamlessly integrate into your organisation’s data strategy and governance framework.

We understand that every organisation is unique. That’s why we don’t offer one-size-fits-all solutions. Instead, we take the time to understand your specific goals, challenges, and data ecosystem. Then, we craft a customised strategy that aligns with your business objectives, ensuring maximum ROI.

Data governance is the cornerstone of effective AI implementation. Our team specialises in developing robust data governance programs that ensure your data is secure, compliant, and ready for AI-driven insights. With our guidance, you can confidently navigate the complex world of data regulations. 

If you’re struggling to envision where Generative AI could fit into your data strategy or are ready to implement a data governance program that sets you up for AI success, have a chat with one of our team.

Zareene Choudhury and Lea Gorgulu Webb are senior consultants here at Oakland

Frequently asked questions

Want to know more about generative AI? Here are answers to some common questions. Don’t see the answer you’re looking for? Get in touch with one of our team.

What is Generative AI?

Generative AI is a type of AI focused on content creation. Gen AI systems are trained on existing data and use it to create original content such as text, images, videos, code, or audio, based on the patterns inherent to the training data. This is different to traditional AI, which only uses the data to make predictions or decisions.

The technology has grown significantly in popularity since November 2022, when ChatGPT was released to the public. While the initial novelty and hype have somewhat subsided, the potential for the technology to drive productivity gains has meant many organisations are adopting generative AI in their operations. A global INSEAD survey of business alumni in 2024 found that just over half of respondents’ organisations were using generative AI and only 21% had no plans to use it in the future.

How Does Generative AI Work?

The workings of generative AI can differ significantly from one use case to another, but in general, they are created using a four-step process that utilises neural networks, and data architectures like transformers:

  1. Data Collection: Significant volumes of data relevant to the desired output are collected. For example, if the goal is to generate customer service responses, transcripts and training materials might be used in the dataset.
  2. Training: The AI model is trained on this dataset using machine learning algorithms. This allows it to notice and replicate patterns in the data.
  3. Generation: Once trained, the model can generate new data that draws on and mimics patterns in the training data.
  4. Fine-Tuning: The model can be fine-tuned for specific applications or industries, ensuring the generated content meets particular standards or requirements.

Is generative AI machine learning?

Generative AI is a system enabled by the broader concept of machine learning. Machine learning algorithms allow generative AI systems to learn from data and then apply these learnings to complete tasks. Generative AI specifically uses those learnings to create original content. 

There are three main ways machine learning algorithms are trained: 

  • Reinforcement learning – Where the machine learning model is trained using rewards and penalties based on its actions. This reward system hones its understanding.
  • Supervised learning – Where data sets are labelled, giving the algorithm insight into the meaning and relationships between different data.
  • Unsupervised learning – Here, the algorithm is given unstructured data without any human input, and allowed to discover relationships and patterns on its own.

Generative AI systems are often developed with the final unsupervised method. This allows them to be more ‘creative’ with their outputs than other machine learning algorithms.

Is generative AI deep learning?

Generative AI is provided using deep learning techniques which involve layered architectures (so-called neural networks) to identify and remember complex patterns in data. Deep learning techniques include things like: 

  • Transformers: Transformers are a neural network architecture that transforms an input sequence (a prompt) into an output sequence (content) based on the learned relationships between the components in the two sequences. They are a key part of large language models like ChatGPT.
  • Generative Adversarial Networks (GANs): This system is formed from a generator, which creates fake outputs that mirror real data, and a discriminator, which tries to guess which are fake or real. The discriminator receives rewards or penalties (via reinforcement learning) if it’s correct or incorrect. GANs can create very lifelike outputs but can be unstable.
  • Variational Autoencoders (VAEs): These unsupervised models understand the structure of data using an encoder, decoder and loss function. The encoder compresses data and gives it a summary (of the image, text, audio, etc.). The system creates a bank of these summaries logged to each piece of compressed data, ready to draw on – red flowers, happy faces, etc.

    The user then provides a prompt, which is given to the decoder. It then tries to reconstruct the data as accurately as possible, based on the summary characteristics. The loss function then measures how much data was lost during reconstruction, distributing data smoothly. In practice, this all means that VAEs can create highly imaginative content (happy red faces, surrounded by petals), but that might be blurry or garbled, due to the loss function. 

The post How to Harness Generative AI for Data-Driven Decision-Making appeared first on Oakland.

]]>
https://weareoakland.com/blog/harnessing-generative-ai-for-data-driven-decision-making/feed/ 0
Why Invest in Data Quality? https://weareoakland.com/blog/why-invest-in-data-quality/ https://weareoakland.com/blog/why-invest-in-data-quality/#respond Wed, 23 Aug 2023 09:31:22 +0000 https://www.theoaklandgroup.co.uk/?p=7551 This can seem like a rhetorical question: you should always invest in Data Quality! But we are still not investing enough: surveys show Data Quality issues are increasing in most organisations and on average, take up 34% of a Data Engineers time instead of them creating value by adding new features. This increases to 50% in large...

The post Why Invest in Data Quality? appeared first on Oakland.

]]>
This can seem like a rhetorical question: you should always invest in Data Quality! But we are still not investing enough: surveys show Data Quality issues are increasing in most organisations and on average, take up 34% of a Data Engineers time instead of them creating value by adding new features. This increases to 50% in large Data Platforms.

All these Data Quality issues add up, with bad Data Quality costing organisations on average $15mil a year.

Data Quality investment is also an investment in high-quality AI and ML, as you’ll likely get more accurate AI and ML results from improving Data Quality than changing your AI model and code.

Having Data Quality checks in place helps reduce “data downtime” for outages and fixes, which subsequently increases the overall reliability of the Data Platform. Highly reliable data leads to more trust in data and better-informed decision-making.

Better decision-making should increase profitability, productivity, and confidence of the whole organisation, which in turn usually leads to more investment in data and, as a result, going back to the start: Increasing the quality of data again.

All this creates a “Virtuous Cycle” of Data Quality, constantly improving your organisation:

If Data Quality decreases the opposite happens with a negative cycle.

Better Data Quality testing should also reduce the blast radius of the issue to a few Data Engineers rather than hundreds or thousands of users as more issues are being found earlier:

The fewer users impacted, the smaller the cost caused by the issue, which should again pay back any investment in Data Quality in large multiples.

Do I need a Data Quality Framework?

I know that setting out to build a framework for anything requires time, and you’ll have many competing concerns, so we understand if you feel reluctant to build one, especially if you are a small team with a limited budget.

But there comes a point where fighting lots of local battles with Data Quality becomes more inefficient than building out a framework to reduce Data Quality issues over the long term.

We’re not going to do a deep dive on Data Quality frameworks here, as they are often tied to wider Data Governance frameworks (you can download our guide here) We will say that whatever framework you use, make sure it’s cyclical so that it’s always improving and you are acting on any emerging issues in a timely manner.

How Do I Test for Data Quality?

Classically, Data Quality tests are a set of rules that test between the actual and desired state of data. The desired state may not be perfect, but ‘good enough’. What counts as ‘good enough‘ varies from dataset to dataset, which makes Data Quality more challenging.

What do we normally test in a dataset, though? The DAMA International’s Guide to the Data Management Body of Knowledge says there are six dimensions to Data Quality:

  • Accuracy: does the date look how we expect it to?
  • Completeness: are there any unexpected missing values?
  • Uniqueness: no duplicates!
  • Consistency: does a person’s data match in two different datasets?
  • Timeliness: is the data out of date?
  • Validity: does the data conform to an expected format? Think postcodes, emails, etc.

Tracking all these dimensions for every dataset at every stage of your pipeline is a lot of work, probably too much work. Therefore, a trade-off is often required to focus on areas where Data Quality will have the most impact on the business.

You also have to beware of false positives or minor issues being blown out of proportion, overwhelming your engineers with too many issues. It can help if you categorise your Data Quality issues by severity just like other software issues.

You also have to take into account the mental wellbeing aspect too: few Data Engineers and Analysts want to spend a large percentage of their time fixing Data Quality issues over a long period of time.

There is help, though, with software frameworks to help you write Data Quality testing:

Most of the above profile your data and setup recommended tests for you to use, saving you some time configuring them yourself.

But in reality, we see a lot of custom-made Data Quality testing, partly because Data Quality struggles for investment, so it is usually done in an organic, ad-hoc manner.

The above products work best in development and staging environments, so you can find issues before they enter production or use them as circuit breakers, to stop a Data Pipeline if the incoming or outgoing data is of poor quality.

It is also worth mentioning that you can use constraints in a Warehouse or Lakehouse schema, which have the benefits of not requiring another software library but are not as feature rich (you will likely have to setup your own notifications for alerting).

It is also important to inform your users of any Data Quality issues as soon as possible so they don’t waste time finding out for themselves or use data that is untrustworthy. This can be done through notifications and alerts, though we’ve also had a lot of success creating Data Quality dashboards that sit alongside existing reports and can be easily referred to by users.

Latest Concepts in Data Quality

There has been significant innovation in Data Quality in the last few years, so we present below the concepts to take your Data Quality process to the next level.

This will require more investment, but it will give you an edge over your competitors to make better informed decisions as you’ll have more trustworthy data. This investment should also pay back long term with less time wasted fixing Data Quality issues.

What is Data Reliability and do I Need it?

Data Reliability gives Data Quality more of a support focus, which makes sense as most Data Quality issues in production will be dealt with as a support issue to a Data Platform.

Data Reliability takes a lot of its thinking from Site Reliability Engineering (SRE), which treats support as more of a engineering problem, where you examine your past and current support tickets and look to decrease them with engineering or better processes.

With Data Reliability, you would look to get a baseline of Data Quality issues per week or month and then look at ways to reduce them and monitor to see if the changes have reduced the number of issues and/or reduced the amount of time spent on issues.

The changes to improve Data Reliability can be technology-based:

  • New or updated tooling
  • Better automation of when a pipeline fails or automated actions to respond to a data issue

Or the changes can be process-oriented:

  • Writing better documentation to avoid common issues
  • Incident playbooks so the whole team can more quickly respond to a issue in an consistent way.

You can rather cynically say Data Reliability is just Data Quality with a feedback loop and a time series graph, but it is there to make sure you avoid short term thinking about Data Quality and instead consider long term improvements that will make your data platform more efficient and trustworthy.

Data Reliability Cycle

You may also set targets such as “99.9% of data will refresh on time” or “A maximum of 33% of engineer time should be spent on support issues“ as well. As mentioned before, it can be impossible to achieve perfect Data Quality, so aiming for a reasonable target instead can avoid engineer burnout.

What is Data Observability and do I Need it?

Data Observability is about gaining a Data Platform or organisation-wide understanding of your Data Quality.

It arguably goes beyond Data Quality by adding metadata features normally found in a Data Catalog: cataloguing schemas of datasets and data lineage. These features allow you to more quickly find a Data Quality issue by tracing the lineage of the issue and also you gain the ability to see how much Data Quality is impacting your organisation.

Data Observability software can often also come with Machine Learning (ML) algorithms to detect anomalies in data, so you can be warned about issues you haven’t even thought of yet.

We’ve seen products either extend a Data Quality framework with Data Catalog features such as Monte Carlo and Big Eye. Or existing Data Catalogs add Data Quality functionality, such as Datahub, which imports Data Quality tests created by Great Expectations and dbt tests. Both Soda and Monte Carlo have integration with the Data Catalog Alation.

What are Data Contracts and do I Need Them?

Data Contracts make a contract between a data producer and a data consumer, so the consumer knows what data to expect from the producer.

While you can replicate some of a Data Contract’s benefits by tracking the schema of the data produced, a Data Contract is meant to go beyond that by giving you a full suite of metadata about the data:

  • The data’s schema.
  • How the data is calculated.
  • Who owns the data?
  • What is the data lineage?
  • How to access the data.
  • What is the data’s expected quality, availability, etc.
  • Plus anything else that is relevant to the data.

You may think Data Contracts are redundant if you have a well-maintained Data Catalog, as they capture similar information, but Data Contracts are designed to be checked during every run of a Data Pipeline and have some action in the pipeline if the Data Contract is broken:

  • Stop the pipeline with a circuit breaker.
  • Alerting.
  • Moving data that doesn’t meet the contract to a manual checking table.

For an example, Paypal has open-sourced their Data Contract template.

Data contract schema

https://github.com/paypal/data-contract-template

This should create more positive collaboration between data producers and consumers because they have a collective agreement of what the data should look like. It is not uncommon to have a poor working relationship where a producer makes changes without telling consumers or consumers accessing data in way not recommended by the producer.

One issue with Data Contracts is that they are a new concept, so require more work to implement at present, though that will likely change in the near future as more companies adopt them.

Most of the examples of Data Contracts we’ve seen so far use Apache Flink and the Kafka Schema Registry, so assume you are using streaming, though there are some examples that use batch processing.

Data Governance and Data Quality

Good Data Governance can also improve quality of data. It is important to know where data is coming from, who owns it, for what purpose data is being transformed, and finally, what is the impact of poor availability and data quality: all helped by having Data Governance properly implemented.

Some of the above concepts (Data Contracts and Data Observability) can also improve Data Governance, so investing in Data Quality can also be an investment in good Governance too.

How Does This All Fit Together?

The diagram below is one example of how it all fits together:

  • Any code changes are tested in development and/or test environments with Data Quality Tests to check that any changes won’t have a negative impact on Data Quality.
  • Source Data at the start of the data pipeline is checked to see if the Data Contract is held; if not, a circuit breaker may kick in, stopping the data pipeline early to avoid processing unsuitable data.
  • Data Quality tests are also run in production, which can feel like duplication from testing in development, but there may be changes caused by moving to a production environment (different data, etc.).
  • Data is collected for observability checks by Data Observability software, looking for any anomalous data: a department budget that goes from £10k to £1mil or 10x increase in rows for a table, for example. This can replace a lot of tests, but not all of them.

You’ll also be collecting Data Quality metadata to improve your Data Reliability.

Making all this work together seamlessly isn’t cheap and will take time, but as mentioned, poor Data Quality will also cost an organisation a lot of money. So we recommend tackling this in an agile manner by improving Data Quality in small increments, one change at a time, starting where it will have the most impact.

Summary

Data Quality is a difficult subject to tackle, due to it being a slightly different problem in every organisation and never “perfect”. That said, there are lots of options to help improve the quality of your data, so you should be able to get to “good enough“ if you give Data Quality enough priority and forethought.

Jake Watson is a Principal Engineer at Oakland

Get In Touch 


The post Why Invest in Data Quality? appeared first on Oakland.

]]>
https://weareoakland.com/blog/why-invest-in-data-quality/feed/ 0
How to create a secure Azure Data Platform https://weareoakland.com/blog/how-to-create-a-secure-azure-data-platform/ https://weareoakland.com/blog/how-to-create-a-secure-azure-data-platform/#respond Tue, 01 Aug 2023 09:34:23 +0000 https://www.theoaklandgroup.co.uk/?p=7524 There are many methods you can use to secure your data platform and the data contained within it within Azure. The security controls that will be most effective for each data platform differ based on the usage of the platform, the data sources for the platform and many other factors; Having a holistic view of...

The post How to create a secure Azure Data Platform appeared first on Oakland.

]]>
There are many methods you can use to secure your data platform and the data contained within it within Azure. The security controls that will be most effective for each data platform differ based on the usage of the platform, the data sources for the platform and many other factors; Having a holistic view of the potential options for securing your platform through each of the security layers below will ensure a platform is both fit for purpose and secure.

Diagram

Azure data platform structure

Implementing and maintaining good security within a data platform can be a timely and expensive endeavour so is often overlooked. It can require specialist knowledge to design and implement and make it more complex to connect systems and resources. However, this has to be balanced against the impact of a security breach which could be financial, reputational and have safety implications. During the design phase of a data platform, the sensitivity of data which it will contain, and potential impacts of a breach should be analysed in order to determine an appropriate level of security controls which should be designed into the platform to appropriately mitigate this risk.

Defence in Depth Security:

  1. Data Governance and Classification
  2. Data Protection
  3. Access Control
  4. Authentication
  5. Network Security
  6. Threat Identification and Remediation
  7. Disaster Recovery

 

  1. Data Governance and Classification

Data Governance and Data classification in the context of security is assigning a security rating to data based on the sensitivity of the information contained. Examples of commonly used classifications are ‘Public’, ‘Internal’ and ‘Confidential’. Data in Azure SQL Databases, Azure SQL Managed Instance and Azure Synapse can be allocated a ‘Classification Label’ and an ‘Information Type’. This can be done manually using T-SQL statements or done within the portal within the ‘Data Discovery & Classification’ tab, which can also automatically infer classifications and information types.

Providing classifications to data is beneficial to security as it enables the ability to monitor access to data of different classifications. Additional data governance tools such as Azure Purview can provide additional data classification features such as the ability to associate classifications to data from sources other than those listed above.

  1. Data Protection

It cannot be assumed that data is protected by default even in Azure Platform as a Service (PaaS) services. The nuances in the differences between data protection between services must be understood in order to create a fully protected data platform. Data encryption is an important part of data protection as it protects data from being useable if it is accessed through malicious activity. Many Azure services provide a certain level of data encryption by default, but this cannot be assumed to be true. Azure SQL Database and Azure SQL Managed Instance both have Transparent Data Encryption (TDE) enabled by default, this service is also available for Azure Synapse Analytics Dedicated SQL Pools, however in this case it is not enabled by default. Given the rise in popularity of lakehouse based architectures within data platforms, it is also worth noting that Azure Data Lake Storage (ADLS) also utilised encryption at rest and in transit, to secure underlying data.

  1. Access Control

Access to data and resources should be granted using the principal of least privilege. So users are only granted access they need to perform their duties and no more. Access should be regularly reviewed and when no longer required it should be revoked. Azure tenants should be carefully designed with Management Groups, Subscriptions and Resource Groups to reduce unnecessary resource visibility, for example preventing the Marketing Department from seeing or accessing Finance Department resources.

Typically the fewer people who have access to a resource or data the more secure it is, this limits the chance of users accidently or deliberately leaking or altering potentially sensitive or business critical data. It’s not only access to data which should be carefully considered, access to resource configuration is also important. For example, if a user is granted contributor access to a resource they could delete or alter the resource by mistake. Additionally, they could make changes which compromise the security of the platform, such as altering or removing Network Security Group (NSG) rules without understanding the consequences, permissions like these should be limited to those with requirement for it and required technical understanding.

Sensitive data can be protected from unauthorised access through the use of data masking. This feature is available to Azure SQL Databases, Azure SQL Managed Instance and Azure Synapse Analytics. The amount of data which can be viewed but different users or user groups can be dynamically defined using policies. For example a Database can be configured so that members of the Finance team are able to view full credit card numbers, HR personnel are able to view the last 4 digits of the card number and developers are only able to a series of ‘X’s. The dynamic data masking feature on the supported services listed above can be configured using T-SQL statements or within the Azure portal on the ‘Dynamic Data Making’ page.

Resource locks can be applied within Azure to control which users can perform certain operations on resources. There are two types of resource locks: ‘Read-Only’ and ‘Delete’. ‘Read-Only’ locks allow users with access to view a resource but they cannot make any changes to it. ‘Delete’ locks stop users from being able to delete resources. These features prevent accidental resource deletion or reconfiguration.

  1. Authentication

Authentication is the method by which users or services verify their identity when attempting to access another service. Azure services have support for authentication and access control with Azure Active Directory (AAD). AAD offers many features to enhance security such as Single Sign On (SSO) which allows users to use one set of credentials to sign on to multiple services.

By default, Role Based Access Controls (RBAC) permissions on resources are granted with authentication using AAD, which can be done for individual users or users grouped into management groups. Data access can also be authenticated using AAD with support for Multi Factor Authentication (MFA) through the use of Azure SQL Database, Azure SQL Managed Instance and Azure Synapse Analytics. This removes the need for additional passwords to access SQL and reduces the likelihood of many users signing in using the admin credentials when this is not required. Also integrating SQL access with AAD means if employees are removed from the organisations AAD for example due to leaving the company then their access to SQL will also be automatically removed without needing to manually delete their SQL user or rotate the password of a shared login.

Authentication is required between services and resources, traditionally this authentication is done using a username and password combination, however this is vulnerable to these credentials being leaked. A more secure method of authenticating between systems is using Managed Identities. When using Managed Identities in Azure, the identity of a resource or service is registered within AAD and other services can use AAD tokens to authenticate, rather than  a username and password.

  1. Network Security

An appropriately designed network security framework in Azure can protect a data platform from attack and unauthorised access, whilst also allowing for functional communication in and out of the network. Features offered by Azure to create a secure and functional network topology include private endpoints to secure PaaS services, Azure firewall and virtual networks with associated NSGs.

Virtual networks create groups of connected services which can be protected from unwanted inbound and outbound communication. Virtual networks are split into subnets, these subnets can have associated NSGs which are lists of rules which allow or deny traffic. When assigning NSGs all communication should be blocked as a default and only exceptions for specific purposes should be made in order to minimise communication. Reducing the number of allowed protocols and allowed ports for communication on your virtual network reduces the routes an attacker could take into your network.

Private endpoints within Azure are available for a range of PaaS services such as Azure Data Lake Gen 2 and Azure SQL Database. A private endpoint creates a Network Interface Card (NIC) which is associated to the virtual network. Then public access to these services can be disabled and only traffic to and from the network can be allowed, creating a private connection. The use of private endpoints can reduce the risk of data leaks as data is always contained within the private network and is not transferred publicly over the internet.

Azure Firewall can be used to inspect and analyse traffic coming from outside an Azure environment into it and between spokes within an Azure hub and spoke model. The firewall can inspect and block unwanted traffic to your Azure environment. Azure Firewall is integrated with Azure Monitor, Azures logging and alerting offering, so metrics from the firewall can be inspected.

Many data platforms require connectivity to on premise systems, these connections can pose a potential security risk of not properly configured and secured. There are several methods for connecting cloud and on premise systems within Azure. When choosing what method to use to connect systems the cost of the connection must be balanced with the required security for the connection. Azure ExpressRoute can provide a dedicated private connection between an Azure environment and an on premise system so no data transferred is exposed to the internet hence making this a very secure option. However Express Route is a costly option. Another option for creating this connection is to use a Site to Site Virtual Private Network (VPN) which transfers encrypted data between the on premise system and Azure over the internet. A Site to Site VPN is typically cheaper than Azure Express Route, but due to data being transferred over the internet it is considered to be less secure.

  1. Threat Identification and Remediation

It can often be difficult to identify security threats to a data platform, you can collect logs and query them in order to identify suspicious or threatening activity but it can often be difficult to filter out unnecessary logs and identify useful information. The Microsoft feature ‘Microsoft Defender for Cloud’ provides detailed recommendations of how resources can be configured to be more secure and of active cyber security alerts such as logins from unusual locations. Defender for Cloud also provides recommendations for remediation steps which should be completed to resolve potential security issues.

Most resources within Azure have the ability to export diagnostic logs and metrics to Azure Monitor or Azure Log Analytics Workspace which are central repositories where logs can be queried and alerts can be set based upon these logs. For example security logs can be exported from Virtual Machines which detail login attempts, then Azure Log Analytics Workspace can be used to query failed login attempts and an alert can be created to email nominated users when a failed login attempt has been made.

  1. Disaster Recovery

In a worst case scenario, a cyber attack (or accidental deletion) could result in services and data being deleted or un recoverable. In this scenario a robust and timely disaster recovery plan which minimises the Recovery Point Objective (RPO) is imperative. This can be achieved through a variety of methods. Storing templates for infrastructure deployment, and ARM templates for orchestration services such as Azure Data Factory in Repos will mean a platform can be redeployed quickly and easily.

Many Azure PaaS storage services have features to simplify recovering lost data which are enabled by default. For example Azure Synapse Analytics creates regular restore points on dedicated SQL pools and a deleted or corrupted database can be redeployed from these restore points.

Azure provides geo-replication features to protect resources and data against a potential disaster at one of its data centres or regions. If enabled on resources geo-replication can create a replica of the resource within a different availability zone or region so if the hardware containing the primary resource is damaged the resource is still available through the replica. This feature is available on many resources in Azure including Azure SQL Database and Azure Storage Accounts.

Abigail is a Senior Data Engineer at Oakland

The post How to create a secure Azure Data Platform appeared first on Oakland.

]]>
https://weareoakland.com/blog/how-to-create-a-secure-azure-data-platform/feed/ 0
How To Create a Data Product-Focused Data Strategy https://weareoakland.com/blog/how-to-create-a-data-product-focused-data-strategy/ https://weareoakland.com/blog/how-to-create-a-data-product-focused-data-strategy/#respond Tue, 02 May 2023 11:51:43 +0000 https://www.theoaklandgroup.co.uk/?p=7246 Data engineering and storage costs are rising rapidly, with a 30% increase in 2022 (US and European cloud prices set to soar in 2023 – Techzine Europe). As US and European cloud prices were set to soar in 2023, data leaders searched for ways to maintain delivery speed while improving data governance and managing budgets....

The post How To Create a Data Product-Focused Data Strategy appeared first on Oakland.

]]>
Data engineering and storage costs are rising rapidly, with a 30% increase in 2022 (US and European cloud prices set to soar in 2023 – Techzine Europe). As US and European cloud prices were set to soar in 2023, data leaders searched for ways to maintain delivery speed while improving data governance and managing budgets. Enter Data Products!

In this blog, we’ll guide you through creating a data product-focused data strategy. We’ll cover key concepts like defining data products, setting up domain-specific teams, and implementing a federated ownership model to enhance data management and efficiency within your organisation.

Are you looking to write a data strategy but unsure how? Read our guide on how to write your data strategy for further information.

What is a Data Product?

Data products are reusable assets that deliver trusted datasets for specific purposes, ensuring the right data reaches the right people in the right format. The unique aspect of data products is the federated ownership model, where domain-specific teams manage their data from production to consumption, ensuring better efficiency and responsibility.

Domain driven ownership

Figure 1 A Domain-Oriented Ownership model

Source: https://www.starburst.io/blog/data-mesh-and-starburst-domain-oriented-ownership-architecture/ 

What is a Data Product-Focused Data Strategy?

A data product-focused data strategy involves creating reusable data assets, known as data products, to solve specific business problems. This strategy enhances data quality, reduces silos, and improves security through an ownership model. 

It shifts the approach from treating data as an output of projects to managing data as products, ensuring ongoing maintenance and updates. This strategy includes defining domains, building data products, and continuously engaging with consumers to improve and deliver valuable data solutions.

Why Do You Need a Data Product-Focused Data Strategy?

In today’s data-centric world, having a solid data product strategy is a necessity for your organisation. 

  1. Improves Data Trustworthiness: Ensures accountability for data quality through domain ownership.
  2. Reduces Data Silos: Facilitates autonomous access to managed data, enhancing self-service capabilities.
  3. Enhances Security: Allows domain teams to manage access and compliance, optimising risk management.
  4. Supports Continuous Improvement: Encourages ongoing maintenance and updates of data assets.
  5. Drives Business Value: Focuses on solving specific business problems, leading to more impactful data solutions.

Read our blog to find out more information about the purpose of a company’s data strategy.

Is the Data Product Approach Suitability for Mature Data Estates?

The data product approach may not suit less mature data estates due to its reliance on advanced processes, skills, and infrastructure. However, adopting it as a long-term goal can improve existing practices and lay the groundwork for future data product delivery.

How to Integrate a Product Data Strategy into your Organisation

Incorporating a data product approach into your data strategy has people, process and technology implications – so the following tips are designed to provide you with some ideas of steps you might want to take if you are considering taking the plunge. 

We mentioned earlier the concept of domain driven ownership. Given the nature of data products being focused on how each domain can solve issues for the business, your first port of call is deciding how to divide the organisation up into domains. 

A good starting point for this is to investigate the existing subject areas that sit within your data estate – most organisations will have these. Once you have a picture of what your domains are, you can start to look at the process of building a data product. 

If your data strategy is led by IT or Data, this is where you start to find friends in the domain areas. Once you have identified a domain that is open to the data product approach, you can work with them to define the data products that would support the business needs from that domain. 

When pitching a product, it must solve a specific problem. Data products are no different. Start by identifying and addressing business challenges.

Key Steps to Build Data Products

You don’t see people on Dragon’s Den pitching a product that doesn’t solve a specific problem. In fact, the most successful companies will think about the problem first and work to build a solution to that problem. Data Products are no different, and building a strategy for these data products should align with that view. 

You will want to build products that solve specific problems and challenges within the business, which can be marketed and maintained. 

Understand Consumer Needs

Get to know your consumers, do your research, understand the pains they feel when engaging with your data. 

Establish Feedback Processes

Establish processes that enable consumers to contact and feedback to data product teams so that every data product is built either toward solving a specific problem or responding to market demand quickly by creating new business opportunities. 

Prioritise

You can then set out a prioritisation exercise to understand where to focus efforts to deliver your first data products! 

Why Should You Take the Lighthouse Approach to Data Strategy?

As with any new process within an organisation, at Oakland – we find that delivering value rapidly is critical. We can do this by taking a lighthouse approach, where you take a narrow but deep slice of your data estate and design a small but valuable data product.

Luke Sharma, Senior Data Consultant at Oakland, states

‘’We achieve this by taking a lighthouse approach, where we focus on a narrow but deep slice of our data estate to design a small yet valuable data product. This method allows us to set up the rituals and roles of product management to fit within our existing operating model. Once the lighthouse project is delivered, we assess the supporting processes and roles, as well as the value gained from the product.’’

He goes on to mention that, ‘’These learnings are then embedded into our Target Operating Model, ensuring the product management approach is integrated into the organisation’s ways of working. The key here is that the approach is tested and drives value for your organisation with a more moderate investment. This lighthouse data product and way of working can then be demonstrated to other domains, increasing appetite and excitement.” 

Breaking the Project Mindset

This is where the strategy element of “data product strategy” becomes vital because it involves cultural and process change. Once you have delivered a data product, you can rapidly see it fall back into the traditional cycle of standard data development because the organisation’s mentality hasn’t moved forward with the technology. 

Data is unfortunately often seen as an output of a project, be it a report, dashboard, or ML algorithm. The project is delivered, signed off by stakeholders and then handed over to the business unit. Inevitably the deliverables from these projects are outdated pretty quickly – and the response? Set up another project, build another report – rinse and repeat. 

You need to stop thinking of data as an output and start to think about data as a product. 

“Products need to be maintained, updated, and associated with SLAs/data contracts that guarantee a certain level of uptime and data quality as part of a standard product lifecycle, managed by the product management team. By taking the approach learned from the lighthouse project and embedding it into a Target Operating Model (TOM), we ensure that accountability for data product maintenance is prioritised.”

Luke Sharma, Senior Data Consultant

The whole aim of data products is they are built and maintained in order to be reusable, shareable assets!

Focusing on Value

Data products must deliver specific outcomes. Engage with consumers to understand their needs and guide product development.

Using data products in your strategy will drive adoption and accountability within your organisation. Engage regularly with consumers to refine and prioritise data products, overcoming technological and process complexities.

For more information, contact us at hello@theoaklandgroup.co.uk. For the original blog post, visit Oakland Blog.

The post How To Create a Data Product-Focused Data Strategy appeared first on Oakland.

]]>
https://weareoakland.com/blog/how-to-create-a-data-product-focused-data-strategy/feed/ 0
How to manage spiralling cloud costs https://weareoakland.com/blog/how-to-manage-spiraling-cloud-costs/ https://weareoakland.com/blog/how-to-manage-spiraling-cloud-costs/#respond Tue, 28 Mar 2023 16:06:01 +0000 https://www.theoaklandgroup.co.uk/?p=7151 “81% of IT teams directed to reduce or halt cloud spending by C-suite” VentureBeats, 2022 “Organisations with little or no cloud cost optimisation plans end up overspending on cloud services by up to 70% without deriving the expected value from it”  Gartner, 2022 Big Tech under pressure from cost-conscious cloud customers (ft.com) The cost of...

The post How to manage spiralling cloud costs appeared first on Oakland.

]]>
“81% of IT teams directed to reduce or halt cloud spending by C-suite”

VentureBeats, 2022

“Organisations with little or no cloud cost optimisation plans end up overspending on cloud services by up to 70% without deriving the expected value from it”

 Gartner, 2022

Big Tech under pressure from cost-conscious cloud customers (ft.com)

The cost of cloud resources has been increasing at an alarming rate for many customers who use cloud service providers such as Microsoft Azure, Amazon Web Services, or Google Cloud, and for those who use multi cloud providers then the costs for cloud usage can be eye-watering. For many, large volumes of data have been integrated into these platforms to capitalise on the huge range of services these cloud providers offer. Ranging from unparalleled data resilience to large-scale advanced analytics.

Initially, these costs may have offered cost savings from the traditional on-premises cost, but as data volumes increase, the cost savings evaporate, and cloud cost management increases with many organisations IT teams overspending on their cloud environment. The good news is there are ways and approaches to reduce cloud costs that can be built into your cloud strategy.

Where do I begin?

At Oakland, we’ve seen a lot of organisations giving more attention to improving Cloud FinOps as a function or approach. Cloud FinOps, short for Cloud Financial Operations, is a set of practices that help optimise cloud spend. Cloud operational management is often decentralised, with costs that can be hard to predict or control. For example, cloud services costs often just arrive weekly or monthly as one large invoice to someone outside of IT with no real transparency on what is being invoiced for. Understanding these costs as an organisation is half the battle.

Following cloud cost management best practices or frameworks can result in the growth and development of cloud solutions in line with manageable costs. Examples of cloud cost management best practices include:

  • Improved collaboration between business and technical stakeholders works towards a common goal of becoming more cost-effective and reducing cloud expenditure.
  • Increase transparency by utilising reporting, resource tagging, and other cost management tools
  • Track and monitor the creation and running of cloud resources.
  • Establish clear internal responsibility of cloud costs to appropriately assign expenses.

What are the other ways to reduce costs?

This is only half the story, as robust Cloud FinOps is underpinned by different processes aiming to improve and enhance an organisations approach to managing cloud spend. There are other ways to reduce costs which range from straightforward to more involved:

Decommissioning

Many organisations have multiple legacy systems storing data. These systems are often:

  • Not used regularly
  • Replaceable with a modern equivalent
  • Expensive to maintain and keep live
  • Require specialist knowledge to integrate with and accommodate

Fewer legacy systems regularly result in simpler data architecture and lower costs from the time invested in the above.

Through identifying valid candidates for decommissioning, a business case can be created to support migrating from and/ or closing down these legacy systems.

Resource planning/scaling

Cloud resources utilise computing power and storage measures to establish an overall cost.

Some processes may be using too many resources or running more often than necessary increasing cloud bills. Reserving resources for longer periods of time will ensure cost optimisation for your cloud expenditure.

Consider ingesting data for a report:

  • Does the ingestion pipeline need to run as often? For how many months/ years?
  • Would it be an issue if it took longer to run?

Amending such factors to scale down resources whilst accounting for the impacts can result in an overall reduction in cloud spending

Cloud Platform Modernisation

Sometimes, a resource-intensive process is still required but costs your IT teams a lot of money to maintain and run.

Re-engineering a pre-existing solution by using more up-to-date and efficient methods may result in cost savings through solutions running faster and being easier for your IT team to maintain.

This approach also allows to alter or augment the solution during this modernisation process.

An up-to-date process that is deemed to be an industry-standard regularly provides more flexibility and cost-saving options than a legacy process or technical approach.

Look at all the money I can save!

Yes, but to a point. Knowing these approaches is different from applying them, as it is often not so straightforward that you can just delete some data, remove a system, shrink resources, or overhaul an inefficient process. Working with a business to know where to practically reduce spending or costs utilises a lot of previous experience rather than turning everything off (or, at worst just doing everything the Azure Advisor says regardless of the impact). Working towards a tangible plan, aligning stakeholders, and initiating the agreed actions are all important requirements to manage any dependencies or concerns and give the business value your organisation is looking for to reduce its IT spending.

At Oakland, we offer a holistic review and outline recommendations and options in an actionable report and give you a deliverable plan. We utilise our vast wealth of experience and are able to look against different lenses to focus the review. These include focusing on architecture design, engineering approaches, governance management, ESG, strategic and tactical alignment, and even market trends. The output is a clear, tailored plan to result in reduced cloud spend / improved Cloud FinOps.

If you would like to find more information about how we can help reduce your cloud spend, please get in touch by emailing hello@theoaklandgroup.co.uk or calling 0113 234 1944.

Jack Evans is a Principal Consultant at the Oakland Group.

The post How to manage spiralling cloud costs appeared first on Oakland.

]]>
https://weareoakland.com/blog/how-to-manage-spiraling-cloud-costs/feed/ 0
What is a data platform? https://weareoakland.com/blog/what-is-a-data-platform/ https://weareoakland.com/blog/what-is-a-data-platform/#respond Thu, 02 Feb 2023 12:34:30 +0000 https://www.theoaklandgroup.co.uk/?p=6904 In today’s digitised business landscape, the ability to collect, analyse, and manage data effectively can be the critical difference between a business’s success and failure. But what exactly enables businesses to harness the full potential of their data? Enter the customer data platform. In this blog, we’ll discuss everything you need to know about data...

The post What is a data platform? appeared first on Oakland.

]]>
In today’s digitised business landscape, the ability to collect, analyse, and manage data effectively can be the critical difference between a business’s success and failure. But what exactly enables businesses to harness the full potential of their data? Enter the customer data platform.

In this blog, we’ll discuss everything you need to know about data platforms, including their many benefits and how they can improve your business. 

What is a Customer Data Platform? 

A customer data platform is a centralised system for collecting, storing, and managing data from various sources. It uses innovative data engineering to leverage components such as data storage, processing, analysis, and integration. This lets organisations gather insights, make data-driven decisions, and ensure data quality and security. 

What is Data Engineering?

Data engineering focuses on the practical aspects of data collection and data analysis. It comprises the design, development, and management of a central system that facilitates the collection, storage, and processing of large volumes of data. 

Key aspects of data engineering include data collection, processing, performance optimisation, security, and seamless collaboration with other teams.

What are the Benefits of Data Engineering?

Data engineering offers a number of handy benefits that can significantly improve a business’s efficiency. The main notable benefits include:

  • Improved Data Quality: Data engineers ensure that the data used for analysis is accurate and reliable through rigorous cleaning and transformation processes.
  • Enhanced Efficiency: Automated data pipelines reduce the time and effort required to manage data, allowing for faster access to insights.
  • Scalability: Data engineering solutions are designed to handle growing volumes of data, ensuring that systems can scale with the business.
  • Better Decision-Making: Data engineering supports data-driven decision-making across the organisation by providing clean, integrated, and accessible data.
  • Cost Savings: Efficient data management and storage solutions can reduce costs associated with data handling and infrastructure.

To learn more about further advantages, visit our dedicated guide: Why do you need a data platform? 

What are Cloud-Native Platforms?

Cloud-native platforms are essential tools to help accelerate the execution of plans over the next 2-3 years. Improved access to cloud services enables the introduction of modern technologies with less operational burden than legacy systems, expediting and facilitating the creation of innovative business solutions.

Adopting cloud-native platforms provides the primary means for enterprises to execute their digital strategies, enabling business growth, customer retention, and efficiency. According to Gartner, cloud-native platforms will serve as the foundation for more than 95% of new digital initiatives by 2025, up from less than 30% in 2021.

What are the Key Components of a Data Platform?

There are several elements that collectively make up a successful digital data platform, including:

Data Access and Governance

In simple terms, a data platform enables data access, governance, delivery, and security. It brings together the technology needed to collect, transform, unify, and govern the data required to support users, applications, models, and data products.

Scalability and Security

To survive in today’s fast-moving market, an organisation’s data platform must be cost-effective, highly scalable, and secure from the outset. 

It should enable data ingestion from multiple sources, including other data platforms, and be flexible enough to accommodate system changes in the future. What’s more, the architecture of your data platform must effectively support your business outcomes.

The Role of Data Governance

A successful data platform combines cloud capabilities with a robust data governance approach. Without data governance, issues like poor data quality and availability can persist, impacting current performance and hindering future growth opportunities.

Targeting Organisational Needs

A data platform must address an organisation’s specific needs. For example, a complex organisation with many data sources will need data conformity as a critical capability. 

We see a data platform as a layered set of capabilities that build on each other. This enables organisations to realise value from high-quality data managed through governance processes, thereby empowering confident decision-making.

The Importance of Cloud-Native Solutions

We believe any modern-day data platform should be cloud-native.

Cloud-native platforms allow organisations to deliver scalable solutions without heavy reliance on managing the underlying infrastructure. These platforms are typically sourced from public cloud services (e.g., Amazon Web Services, Microsoft Azure, Google Cloud Platform) or created using software that structures a private cloud environment for added security and control.

What are the Core Capabilities of Cloud-Native Platforms?

Cloud-native platforms use core functionalities such as container management, infrastructure-as-code, and serverless functions while supporting continuous integration and delivery pipelines. These platforms can work with other cloud tools, SaaS tools, or on-premise applications, offering a speedier alternative to traditional on-premise solutions.

Enabling core capabilities (shown in the diagram below under the themes of knowledge, insight, and awareness) can help realise the true value of the cloud through the provision of all three layers— infrastructure, services, and governance.

As you seek to progress from knowledge to insight and awareness, your capabilities must evolve to meet your ambitions. The component parts of a data platform are shown below.

Data Journey 

How to Build a Data Platform

The easiest way to visualise building a data platform is to think of it in terms of building a house. 

When building a house (data platform), you don’t just start laying bricks (processing data). You need to know the room measurements (data subject areas), layout (data models), and adherence to building regulations (governance). This is how you design a data platform that delivers on your business goals.

For more expert knowledge on building a data platform, visit our blog: What are the challenges of building a data platform?

What are the Benefits of Cloud-Native Platforms?

As technology advances, we’ve seen a number of evolving benefits to opting for a Cloud-Native platform for your data strategy. These include: 

  • Reduced Total Cost of Ownership: Cloud data platform deployment and operation costs are generally lower than traditional on-premise implementations.
  • Service Resiliency and Management: Cloud data platforms require less effort and complexity to meet spikes in demand while maintaining operations.
  • Speed and Service Agility: Cloud-based platforms significantly reduce delivery timescales, leveraging reusable cloud blueprint architectures and components.
  • Business Model Transformation/Optimisation: Cloud-native platforms support broader benefits beyond cost and speed, enabling business model transformation and optimisation.

With O’Reilly survey data showing that 77% of organisations were already cloud-native or pursuing a cloud-first strategy in 2021, the benefits are already being understood across industries.

Discover More Data Strategy Solutions with Oakland

Ready to transform your business with a modern cloud-based platform? Download our guide to delivering one, or contact us if you have any questions.

Want to stay up to date with the latest industry insights on all things data? Explore the Oakland blog today.

The post What is a data platform? appeared first on Oakland.

]]>
https://weareoakland.com/blog/what-is-a-data-platform/feed/ 0
Does Microsoft Purview solve the Data Governance Challenge? https://weareoakland.com/blog/does-microsoft-purview-solve-the-data-governance-challenge/ https://weareoakland.com/blog/does-microsoft-purview-solve-the-data-governance-challenge/#respond Tue, 22 Nov 2022 13:11:36 +0000 https://www.theoaklandgroup.co.uk/?p=6855 Microsoft Purview Data Governance for the cloud, on-premise, multi-cloud and office 365 workloads. Introduction Most organisations are exploding with data that has been collected, transformed, and reported on, but this data is often not well-tracked as the organisation becomes more data-driven, increasing two pain problems that have been growing for the last few decades: How...

The post Does Microsoft Purview solve the Data Governance Challenge? appeared first on Oakland.

]]>
Microsoft Purview

Data Governance for the cloud, on-premise, multi-cloud and office 365 workloads.

Introduction

Most organisations are exploding with data that has been collected, transformed, and reported on, but this data is often not well-tracked as the organisation becomes more data-driven, increasing two pain problems that have been growing for the last few decades:

  • How can we audit all this data to protect against data leaks and unexpected data loss?
  • How can data users discover data in an environment that changes constantly?

Data Governance Products help mitigate these problems, among others, but are often complex due to requiring:

  • The ability to scan a large variety of data sources
  • A highly customised user interface
  • A powerful search engine to find data assets by many different types of metadata attributes

These are just some of the main requirements that create a software marketplace full of products that are often expensive and hard to implement and maintain.

These products also need to ingest large amounts of sensitive organisational data to meet user requirements, ironically creating a Data Governance concern in itself!

Purview aims to ease the pain of Data Governance by being feature-rich, easy to deploy, maintain and secure. But is it worth the cost, and can it compete with bespoke Data Governance companies that have a head start measured in years or even decades?

Purviews Features

  • Its connectors are very Microsoft-focused but cover most of its ecosystem: Azure, SQL Server, Power BI, and Office 365. If you’ve already bought heavily into Microsoft, you can scan most or all your data assets automatically.
  • It focuses less on connectors made by other companies but still covers many popular data products like SAP, Salesforce, Oracle, GCP Big Query, AWS S3, and Snowflake.
  • It offers a lot of flexibility in managing Data Catalog users with 9 different roles to choose from and syncs up to your Azure Active Directory groups and users.
  • Can classify data with 200+ pre-built classifications, as well as custom classifications.
  • Business Glossary with an extensive text editor. Ability to add contacts for roles like Data Owner and Steward to each data asset.
  • Pre-built reports to quickly check insights such as what percentage of data has a Data Owner and the percentage of new data assets in the last month.
  • Offers an API and Python SDK for making custom data sources where connectors don’t exist or mass updating existing scanned data assets.
  • It doesn’t offer much insight into Data Quality of scanned data assets, which can be found in other Data Governance products. However, it could theoretically push Data Quality metrics to Purview via its API.
  • Data Sharing allows users to give other users read-only Data Lake data access without having to copy data.
  • Can ingrate Data Governance with Master Data Management using Profisee
  • Purview is relatively new after only being available to customers for a few years but it is receiving heavy investment from Microsoft, with new features appearing monthly.

Deployment

  • As someone who has designed and built many data platforms, I highly value any product that can be deployed quickly, has low maintenance, and will meet strong client IT & security requirements. I believe Purview is stronger than most Data Governance products in this area.
  • It is as easy to deploy and maintain in Azure as any SaaS data governance product but also offers a choice – 20 plus regions to deploy into, including the UK.
  • Purview can also keep all traffic in and out of its server on its private network using Private Endpoints, never touching the public internet, offering an extra layer of data security when creating a Data Catalogue.
  • Can scan Azure Data Products via Managed Identity authentication offering high-security data connections without worrying about managing passwords.
  • It can connect directly to scan on-premise and other cloud data assets, though it requires some technical knowledge.

Costs

Automated Data Governance tooling is often not cheap, with costs starting in the thousands of pounds for most products. Purview is no exception: it has a base price of £250 per month and will cost more if scanning large workloads. However extra capacity is costed using the pay as you go model, so you only get charged extra when scanning lots of data.

There are also additional extra costs depending on which features are used.

Due to the pricing being highly variable in Purview, many organisations will build a Proof of Concept to road-test Purview for a month or so to accurately measure costs.

Alternatives

Note this isn’t a comprehensive list and is a quickly evolving space with new exciting start-ups entering all the time.

  • Build your own:
    • Excel: low cost, low maintenance if data structures don’t update regularly, doesn’t require any specialist skills to build. While we suspect this is the most common type of data catalogue used, we feel nervous about doing a Data Catalog in a data tool infamous for having poor Data Governance. It does not scale and requires lots of manual effort to work out data lineages and classify data sensitivity.
    • Automate your own solution by extracting schemas of databases and files. This is a nice quick way of generating a data catalogue with low maintenance and little extra costs. You can also build a dashboard on top of the Business Intelligence (BI) platform of your choice. It requires minimum effort if the number of data assets is small, though adding features like data lineage and classifying data sensitivity will require a reasonable amount of engineering effort.
  • Databricks Unity Catalog – ideal for Databricks heavy data platforms, as it is a free extra. Though it will only scan what Databricks can scan. You can integrate with other Data Governance products and update schema as they update in real time, which you don’t see much in other Data Governance products.
  • Other clouds Data Governance solutions like AWS Glue Data Catalog and GCP Dataplex. Both are arguably less feature-rich than Purview, especially for non-technical users, though they are easier to implement if most of your data assets are in their respective clouds. Also, both are used to ingest into other larger Data Governance products.
  • Mature products like Informatica and Talend. These tend to charge by the user and are more commonly found on-premise (though they can be configured and maintained in the cloud on Virtual Machines). They will likely cost the most; sometimes, this is significant, but these are well-trusted and reliable. Add the most value if you buy into the rest of their data platform ecosystems.
  • New products like Atlan and Immuta often focus on providing Data Governance to more recent data products like Databricks and Snowflake but also often focus on making deployments into the cloud more accessible by offering deployments via Docker or Kubernetes. Immuta also provides a single pane of glass for fine-grain data access across many popular data products that allows data access controls at a column and row level.
  • Open Source software like Datahub and Amundsen, both built by large tech companies (LinkedIn and Lfyt respectively). These are the go solutions if your organisation has the technical capacity to build and maintain complex workflows. They offer a lot of customisation and no licence costs, so they can be much cheaper at scale and be more custom tailored to fit an organisation’s Data Governance needs.

Summary

In an increasingly challenging data governance market, Azure Purview is a serious option to consider despite missing some features compared to more bespoke Data Governance companies.

However, using Purview if you spend hundreds per month or more on Data Governance is not a good return on investment, or you want a solution that easily integrates Data Governance and Quality.

If you are looking for a Data Governance product that is easy to deploy, secure, catalogue, and classify data assets, and provides some customisation through APIs and user interface at a competitive cost, then we think Purview is a good contender.

Jake Watson is a Senior Data Engineer at The Oakland Group

The post Does Microsoft Purview solve the Data Governance Challenge? appeared first on Oakland.

]]>
https://weareoakland.com/blog/does-microsoft-purview-solve-the-data-governance-challenge/feed/ 0