{"id":4694,"date":"2020-03-01T15:13:02","date_gmt":"2020-03-01T15:13:02","guid":{"rendered":"https:\/\/www.theoaklandgroup.co.uk\/?p=4694"},"modified":"2020-03-01T15:13:02","modified_gmt":"2020-03-01T15:13:02","slug":"why-data-quality-matters","status":"publish","type":"post","link":"https:\/\/weareoakland.com\/blog\/why-data-quality-matters\/","title":{"rendered":"Why data quality matters"},"content":{"rendered":"<p><strong>Introduction &#8211; Why Data Quality matters and the Great Expectations library<\/strong><\/p>\n<p>Over the last decade or so companies have been striving to make better use of their data. The use cases for such projects have generally fallen under two categories, improve operational efficiency or drive customer sales\/behaviour. However, in order to utilise this data it must first be piped from source systems (CRM, ordering, POS etc) into somewhere with greater redundancy. In addition to this simpler goal, it must also be manipulated into a format that is acceptable to the people analysing said data!<\/p>\n<p>The pipelines that complete these operations are often complex, incorporating numerous source systems with different schemas, update times etc. This leads to the development of numerous functions that work within an ETL flow, scheduled by a tool such as\u00a0<a href=\"https:\/\/docs.prefect.io\/core\/\"><span>Prefect<\/span><\/a>\u00a0or\u00a0<a href=\"https:\/\/airflow.apache.org\/\"><span>Airflow<\/span><\/a>.<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-4769 aligncenter\" src=\"https:\/\/weareoakland.com\/wp-content\/uploads\/2024\/01\/etl_tools.jpg\" alt=\"\" width=\"492\" height=\"272\" \/><\/p>\n<p>&nbsp;<\/p>\n<p>So with the crucial nature of this data in mind, how do we ensure that what gets pulled though our flow is going to be of use to those at the end of the pipeline? That something hasn\u2019t been misentered or corrupted in the source systems? Well in most software the role of unit\/integration testing would help, however, if your unit test expects a data frame and to return a dataframe, as a simple example, this may pass whilst the data within said dataframe is riddled with NULL values and bad quality data. As such, a good addition to more standard testing is to actually test what data those source systems are providing, such as in the example below, which is where a fairly recent library,\u00a0<a href=\"https:\/\/github.com\/great-expectations\"><span>Great Expectations<\/span><\/a>\u00a0(GE), can be a real help! (Though as is usually the case, other DQ testing libraries exist that can be explored, such as\u00a0<a href=\"https:\/\/github.com\/awslabs\/deequ\"><span>deequ<\/span><\/a>\u00a0and\u00a0<a href=\"https:\/\/github.com\/ZaxR\/bulwark\"><span>bulwark<\/span><\/a>)<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-4772 aligncenter\" src=\"https:\/\/weareoakland.com\/wp-content\/uploads\/2024\/01\/pipeline_testing.jpg\" alt=\"\" width=\"652\" height=\"497\" \/><\/p>\n<p>&nbsp;<\/p>\n<p>To summarise the library, GE works in addition to unit\/integration testing by profiling data sources in order to build a set of \u201cexpectations\u201d around each column based on the type, which can then be pruned\/added to by the user. For example, on the dataframe shown below we have an ID column that you may expect never to be NULL, an Animal column that should always be a string and a cost that should always be a float.<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-4775 aligncenter\" src=\"https:\/\/weareoakland.com\/wp-content\/uploads\/2024\/01\/pd.jpg\" alt=\"\" width=\"753\" height=\"288\" \/><\/p>\n<p>An example flow of a GE pipeline would involve initial profiling of data, adaption of the produced expectations and continual validation of new data against said expectations. In addition to a better understanding of the source and processed data, this process can help visualise it though auto-generated html, which can be deployed as a static website. This nice addition allows the user to deploy a very quick data quality dashboard, an example of which is shown below.<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-4765 aligncenter\" src=\"https:\/\/weareoakland.com\/wp-content\/uploads\/2024\/01\/data_docs.jpg\" alt=\"\" width=\"778\" height=\"715\" \/><\/p>\n<p><strong>Getting started with GE on Databricks<\/strong><\/p>\n<p>As eluded to in the title of this post, we have been utilising GE on the well-known Spark platform Databricks, using this platform across a number of clients in order to do a distributed computation of large datasets. However, whilst this post is based around Spark, GE can work with other datatypes such as CSVs (via Pandas Dataframes) and Relational Databases (via SQL Alchemy).<\/p>\n<p>So, in order to start using GE in Databricks you must first follow the initialisation\u00a0instructions on your local machine, after downloading from PyPy both locally and on your Databricks cluster. During set up choose option 1 regarding data sources and then 2 for pyspark, which will give you an error unless you have pyspark installed locally, however this doesn&#8217;t matter. If you now check the directory you initiated within you&#8217;ll find a great_expectations folder and a great_expectations.yml file which you can open and edit. In this file add the following code, so that it matches the image shown below.<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-4773 aligncenter\" src=\"https:\/\/weareoakland.com\/wp-content\/uploads\/2024\/01\/yaml.jpg\" alt=\"\" width=\"844\" height=\"186\" \/><\/p>\n<p>Once this is saved you can copy the local files over into the DataBricks File System (DBFS), using the DataBricks CLI, as shown below and documented\u00a0<a href=\"https:\/\/docs.databricks.com\/dev-tools\/cli\/dbfs-cli.html\"><span>here<\/span><\/a>. This involves first making a directory in the DBFS and then copying over the files.<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-4774 aligncenter\" src=\"https:\/\/weareoakland.com\/wp-content\/uploads\/2024\/01\/mkdir.jpg\" alt=\"\" width=\"758\" height=\"98\" \/><\/p>\n<p>Once this is done you can check the files have copied by opening a notebook, and using the %fs cell magic to check the contents of your dbfs, using the below code.<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-4781 aligncenter\" src=\"https:\/\/weareoakland.com\/wp-content\/uploads\/2024\/01\/fs.jpg\" alt=\"\" width=\"758\" height=\"81\" \/><\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-4771 aligncenter\" src=\"https:\/\/weareoakland.com\/wp-content\/uploads\/2024\/01\/ge_dbfs.jpg\" alt=\"\" width=\"518\" height=\"185\" \/><\/p>\n<p>Now that you have the initilisation set up and the library on your cluster you can set your GE context (which tells the profiler what kind of data to expect) and set your initial expectations, which we&#8217;ll document in two different ways, manual profiling and full profiling, the latter involving the building of the above-mentioned data docs.<\/p>\n<p><strong>Setting your expectations &#8211; Setup<\/strong><\/p>\n<p>To do either of the above-mentioned methods there is a standard set up, which is shown below. This sets up your data context and builds a list of your data assets in a specific database, and then adds a list with paths of where to store the expectations.<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-4780 aligncenter\" src=\"https:\/\/weareoakland.com\/wp-content\/uploads\/2024\/01\/op1.jpg\" alt=\"\" width=\"758\" height=\"407\" \/><\/p>\n<p><strong>Method 1: Manual Profiling<\/strong><\/p>\n<p>The first step to manual profiling is to load the required data, convert this to a GE recognisable spark dataframe object and then create an empty expectations file.<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-4779 aligncenter\" src=\"https:\/\/weareoakland.com\/wp-content\/uploads\/2024\/01\/op1c.jpg\" alt=\"\" width=\"759\" height=\"152\" \/><\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-4766 aligncenter\" src=\"https:\/\/weareoakland.com\/wp-content\/uploads\/2024\/01\/empty_expectation.jpg\" alt=\"\" width=\"533\" height=\"96\" \/><\/p>\n<p>Once you have your empty file you can manually add in the relevant expectations for the desired columns, and profile the batch of data via the validation operator and save out the new expectation suite (in this scenario also saving any failed expectations).<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-4778 aligncenter\" src=\"https:\/\/weareoakland.com\/wp-content\/uploads\/2024\/01\/op1cv.jpg\" alt=\"\" width=\"755\" height=\"151\" \/><\/p>\n<p><strong>Method 2: Full Profiling<\/strong><\/p>\n<p>As opposed to manual profiling, full profiling utilises the inbuilt data profiler that is packaged with GE in order to build a full expectation suite based on the column types, and then validate the data against all of these expectations. A workflow for such a method might be to profile all the tables within a database, maintain these expectations and then validate any updates from the raw data that is appended to these tables. This example is what is shown below, which utilises a couple of custom functions which can be applied to each table within the data asset list created during set up.<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-4777 aligncenter\" src=\"https:\/\/weareoakland.com\/wp-content\/uploads\/2024\/01\/op2c.jpg\" alt=\"\" width=\"754\" height=\"708\" \/><\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-4776 aligncenter\" src=\"https:\/\/weareoakland.com\/wp-content\/uploads\/2024\/01\/op2cr.jpg\" alt=\"\" width=\"754\" height=\"225\" \/><\/p>\n<p>The first function should produce an expectations file for each table, which can then be validated against to build a new data docs page. Examples of the json outputs from both functions are shown below.<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-4767 aligncenter\" src=\"https:\/\/weareoakland.com\/wp-content\/uploads\/2024\/01\/empty_profile.jpg\" alt=\"\" width=\"1005\" height=\"85\" \/><\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-4770\" src=\"https:\/\/weareoakland.com\/wp-content\/uploads\/2024\/01\/full_profile.jpg\" alt=\"\" width=\"1790\" height=\"164\" \/><\/p>\n<p><strong>Wrap up<\/strong><\/p>\n<p>So, hopefully this post has shown you a potential new way to more simply test the quality of what\u2019s moving within your data pipeline, rather than just the functions that make it up! In addition to the enhanced testing that GE provides, We believe that the incorporation of the static Data Docs data quality pages could be a real help to organisations looking to quickly understand the quality<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Introduction &#8211; Why Data Quality matters and the Great Expectations library Over the last decade or so companies have been striving to make better use of their data. The use cases for such projects have generally fallen under two categories, improve operational efficiency or drive customer sales\/behaviour. However, in order to utilise this data it&#8230;<\/p>\n","protected":false},"author":3,"featured_media":3887,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"content-type":"","footnotes":""},"categories":[160],"tags":[74,8,126,161,162,163,145],"class_list":["post-4694","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-tech-talk","tag-data-quality","tag-data-science","tag-databricks","tag-ge","tag-great-expectations","tag-rich-louden","tag-spark"],"acf":[],"yoast_head":"<!-- This site is optimized with the Yoast SEO Premium plugin v27.2 (Yoast SEO v27.2) - https:\/\/yoast.com\/product\/yoast-seo-premium-wordpress\/ -->\n<title>Why data quality matters | Oakland<\/title>\n<meta name=\"description\" content=\"Using Databricks to test the quality of what\u2019s moving within your data pipeline\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/weareoakland.com\/blog\/why-data-quality-matters\/\" \/>\n<meta property=\"og:locale\" content=\"en_GB\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Why data quality matters\" \/>\n<meta property=\"og:description\" content=\"Using Databricks to test the quality of what\u2019s moving within your data pipeline\" \/>\n<meta property=\"og:url\" content=\"https:\/\/weareoakland.com\/blog\/why-data-quality-matters\/\" \/>\n<meta property=\"og:site_name\" content=\"Oakland\" \/>\n<meta property=\"article:published_time\" content=\"2020-03-01T15:13:02+00:00\" \/>\n<meta name=\"author\" content=\"NicolaThomsonOKG\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:creator\" content=\"@dataispeople\" \/>\n<meta name=\"twitter:site\" content=\"@dataispeople\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"NicolaThomsonOKG\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"8 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\/\/weareoakland.com\/blog\/why-data-quality-matters\/#article\",\"isPartOf\":{\"@id\":\"https:\/\/weareoakland.com\/blog\/why-data-quality-matters\/\"},\"author\":{\"name\":\"NicolaThomsonOKG\",\"@id\":\"https:\/\/weareoakland.com\/#\/schema\/person\/2c1329190831e35a1d7a01b1bfacd90a\"},\"headline\":\"Why data quality matters\",\"datePublished\":\"2020-03-01T15:13:02+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\/\/weareoakland.com\/blog\/why-data-quality-matters\/\"},\"wordCount\":1100,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\/\/weareoakland.com\/#organization\"},\"image\":{\"@id\":\"https:\/\/weareoakland.com\/blog\/why-data-quality-matters\/#primaryimage\"},\"thumbnailUrl\":\"\",\"keywords\":[\"data quality\",\"data science\",\"Databricks\",\"GE\",\"great expectations\",\"rich louden\",\"Spark\"],\"articleSection\":[\"Tech Talk\"],\"inLanguage\":\"en-GB\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\/\/weareoakland.com\/blog\/why-data-quality-matters\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\/\/weareoakland.com\/blog\/why-data-quality-matters\/\",\"url\":\"https:\/\/weareoakland.com\/blog\/why-data-quality-matters\/\",\"name\":\"Why data quality matters | Oakland\",\"isPartOf\":{\"@id\":\"https:\/\/weareoakland.com\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\/\/weareoakland.com\/blog\/why-data-quality-matters\/#primaryimage\"},\"image\":{\"@id\":\"https:\/\/weareoakland.com\/blog\/why-data-quality-matters\/#primaryimage\"},\"thumbnailUrl\":\"\",\"datePublished\":\"2020-03-01T15:13:02+00:00\",\"description\":\"Using Databricks to test the quality of what\u2019s moving within your data pipeline\",\"breadcrumb\":{\"@id\":\"https:\/\/weareoakland.com\/blog\/why-data-quality-matters\/#breadcrumb\"},\"inLanguage\":\"en-GB\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/weareoakland.com\/blog\/why-data-quality-matters\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-GB\",\"@id\":\"https:\/\/weareoakland.com\/blog\/why-data-quality-matters\/#primaryimage\",\"url\":\"\",\"contentUrl\":\"\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/weareoakland.com\/blog\/why-data-quality-matters\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/weareoakland.com\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Why data quality matters\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/weareoakland.com\/#website\",\"url\":\"https:\/\/weareoakland.com\/\",\"name\":\"Oakland\",\"description\":\"\",\"publisher\":{\"@id\":\"https:\/\/weareoakland.com\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/weareoakland.com\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-GB\"},{\"@type\":\"Organization\",\"@id\":\"https:\/\/weareoakland.com\/#organization\",\"name\":\"Oakland\",\"url\":\"https:\/\/weareoakland.com\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-GB\",\"@id\":\"https:\/\/weareoakland.com\/#\/schema\/logo\/image\/\",\"url\":\"https:\/\/weareoakland.com\/wp-content\/uploads\/2024\/02\/Oakland-White-logo.png\",\"contentUrl\":\"https:\/\/weareoakland.com\/wp-content\/uploads\/2024\/02\/Oakland-White-logo.png\",\"width\":3372,\"height\":648,\"caption\":\"Oakland\"},\"image\":{\"@id\":\"https:\/\/weareoakland.com\/#\/schema\/logo\/image\/\"},\"sameAs\":[\"https:\/\/x.com\/dataispeople\",\"https:\/\/www.linkedin.com\/company\/oakland-group\"]},{\"@type\":\"Person\",\"@id\":\"https:\/\/weareoakland.com\/#\/schema\/person\/2c1329190831e35a1d7a01b1bfacd90a\",\"name\":\"NicolaThomsonOKG\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-GB\",\"@id\":\"https:\/\/secure.gravatar.com\/avatar\/?s=96&d=mm&r=g\",\"url\":\"https:\/\/secure.gravatar.com\/avatar\/?s=96&d=mm&r=g\",\"contentUrl\":\"https:\/\/secure.gravatar.com\/avatar\/?s=96&d=mm&r=g\",\"caption\":\"NicolaThomsonOKG\"},\"url\":\"https:\/\/weareoakland.com\/blog\/author\/nicolathomsonokg\/\"}]}<\/script>\n<!-- \/ Yoast SEO Premium plugin. -->","yoast_head_json":{"title":"Why data quality matters | Oakland","description":"Using Databricks to test the quality of what\u2019s moving within your data pipeline","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/weareoakland.com\/blog\/why-data-quality-matters\/","og_locale":"en_GB","og_type":"article","og_title":"Why data quality matters","og_description":"Using Databricks to test the quality of what\u2019s moving within your data pipeline","og_url":"https:\/\/weareoakland.com\/blog\/why-data-quality-matters\/","og_site_name":"Oakland","article_published_time":"2020-03-01T15:13:02+00:00","author":"NicolaThomsonOKG","twitter_card":"summary_large_image","twitter_creator":"@dataispeople","twitter_site":"@dataispeople","twitter_misc":{"Written by":"NicolaThomsonOKG","Est. reading time":"8 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/weareoakland.com\/blog\/why-data-quality-matters\/#article","isPartOf":{"@id":"https:\/\/weareoakland.com\/blog\/why-data-quality-matters\/"},"author":{"name":"NicolaThomsonOKG","@id":"https:\/\/weareoakland.com\/#\/schema\/person\/2c1329190831e35a1d7a01b1bfacd90a"},"headline":"Why data quality matters","datePublished":"2020-03-01T15:13:02+00:00","mainEntityOfPage":{"@id":"https:\/\/weareoakland.com\/blog\/why-data-quality-matters\/"},"wordCount":1100,"commentCount":0,"publisher":{"@id":"https:\/\/weareoakland.com\/#organization"},"image":{"@id":"https:\/\/weareoakland.com\/blog\/why-data-quality-matters\/#primaryimage"},"thumbnailUrl":"","keywords":["data quality","data science","Databricks","GE","great expectations","rich louden","Spark"],"articleSection":["Tech Talk"],"inLanguage":"en-GB","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/weareoakland.com\/blog\/why-data-quality-matters\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/weareoakland.com\/blog\/why-data-quality-matters\/","url":"https:\/\/weareoakland.com\/blog\/why-data-quality-matters\/","name":"Why data quality matters | Oakland","isPartOf":{"@id":"https:\/\/weareoakland.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/weareoakland.com\/blog\/why-data-quality-matters\/#primaryimage"},"image":{"@id":"https:\/\/weareoakland.com\/blog\/why-data-quality-matters\/#primaryimage"},"thumbnailUrl":"","datePublished":"2020-03-01T15:13:02+00:00","description":"Using Databricks to test the quality of what\u2019s moving within your data pipeline","breadcrumb":{"@id":"https:\/\/weareoakland.com\/blog\/why-data-quality-matters\/#breadcrumb"},"inLanguage":"en-GB","potentialAction":[{"@type":"ReadAction","target":["https:\/\/weareoakland.com\/blog\/why-data-quality-matters\/"]}]},{"@type":"ImageObject","inLanguage":"en-GB","@id":"https:\/\/weareoakland.com\/blog\/why-data-quality-matters\/#primaryimage","url":"","contentUrl":""},{"@type":"BreadcrumbList","@id":"https:\/\/weareoakland.com\/blog\/why-data-quality-matters\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/weareoakland.com\/"},{"@type":"ListItem","position":2,"name":"Why data quality matters"}]},{"@type":"WebSite","@id":"https:\/\/weareoakland.com\/#website","url":"https:\/\/weareoakland.com\/","name":"Oakland","description":"","publisher":{"@id":"https:\/\/weareoakland.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/weareoakland.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-GB"},{"@type":"Organization","@id":"https:\/\/weareoakland.com\/#organization","name":"Oakland","url":"https:\/\/weareoakland.com\/","logo":{"@type":"ImageObject","inLanguage":"en-GB","@id":"https:\/\/weareoakland.com\/#\/schema\/logo\/image\/","url":"https:\/\/weareoakland.com\/wp-content\/uploads\/2024\/02\/Oakland-White-logo.png","contentUrl":"https:\/\/weareoakland.com\/wp-content\/uploads\/2024\/02\/Oakland-White-logo.png","width":3372,"height":648,"caption":"Oakland"},"image":{"@id":"https:\/\/weareoakland.com\/#\/schema\/logo\/image\/"},"sameAs":["https:\/\/x.com\/dataispeople","https:\/\/www.linkedin.com\/company\/oakland-group"]},{"@type":"Person","@id":"https:\/\/weareoakland.com\/#\/schema\/person\/2c1329190831e35a1d7a01b1bfacd90a","name":"NicolaThomsonOKG","image":{"@type":"ImageObject","inLanguage":"en-GB","@id":"https:\/\/secure.gravatar.com\/avatar\/?s=96&d=mm&r=g","url":"https:\/\/secure.gravatar.com\/avatar\/?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/?s=96&d=mm&r=g","caption":"NicolaThomsonOKG"},"url":"https:\/\/weareoakland.com\/blog\/author\/nicolathomsonokg\/"}]}},"_links":{"self":[{"href":"https:\/\/weareoakland.com\/wp-json\/wp\/v2\/posts\/4694","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/weareoakland.com\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/weareoakland.com\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/weareoakland.com\/wp-json\/wp\/v2\/users\/3"}],"replies":[{"embeddable":true,"href":"https:\/\/weareoakland.com\/wp-json\/wp\/v2\/comments?post=4694"}],"version-history":[{"count":0,"href":"https:\/\/weareoakland.com\/wp-json\/wp\/v2\/posts\/4694\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/weareoakland.com\/wp-json\/"}],"wp:attachment":[{"href":"https:\/\/weareoakland.com\/wp-json\/wp\/v2\/media?parent=4694"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/weareoakland.com\/wp-json\/wp\/v2\/categories?post=4694"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/weareoakland.com\/wp-json\/wp\/v2\/tags?post=4694"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}