The Enterprise Data Reference Architecture (EDRA) serves as a logical progression following the crafting of the data strategy and definition of the enterprise data framework. Essentially, it represents the envisioned state of the enterprise data landscape, reflecting the desired data management and information usage.
While the enterprise data framework delineates the processes and capabilities for data management, the EDRA goes further by illustrating how these principles are organised to facilitate the delivery of enterprise data. Moreover, it guides the design that will provide efficient data management and agile access to information.
The EDRA relies on standardised terminology to efficiently classify and group technology components. This standardisation ensures alignment across the enterprise, streamlining different initiatives, and facilitating communication and collaboration among stakeholders. Additionally, the groupings may intersect with other Enterprise Architecture (EA) domains. Identifying common technology early in the process is crucial for streamlining the EA strategy, enabling effective alignment of resources and priorities. For example, recognising that data ingestion processes overlap with Applications Architecture and Infrastructure Architecture allows for integrated planning and implementation, maximising efficiency and minimising duplication of effort.
If known, each grouping within the EDRA should include elements that serve as starting points for discussion. For instance, consider the element of real-time data ingestion: if sensors are already installed on equipment or external locations, existing processes likely exist for ingesting data. This enables not only identification of the pattern elements related to real-time data ingestion but also their mapping to specific strategic technologies.
Conversely, if data protection is a new discipline for an organisation, the creation of elements and their mapping to specific technologies requires a thorough understanding of business needs, security architecture standards, and regulatory requirements. In cases where these details are unknown, placeholders can be created without specific components, or by including all available methods and technology options. For instance, data protection measures may encompass secret zones, encryption, masking, or hash-keying techniques, which can be explored and selected at a later stage.
Now that we’ve covered the conceptual aspects of the EDRA, let’s explore how they are put into practice. The simple diagram below depicts the architecture with only main groupings, also demonstrates interdependence with other EA domains, where Infrastructure would include Security Architecture as well.

To illustrate the next level of the EDRA, let’s draw parallels with water resources and consumption. Just as water originates from various sources such as rivers, lakes, wells, and sea and undergoes different processes for purification and desalination, data is ingested by various methods such as batch processes, manual entry, messaging, and streaming. Building upon the earlier example, sensors data may be ingested using messaging or streaming patterns, while customer details updates could occur through direct input via dedicated CRM forms or be ingested via batch processes from a CRM system. There will be a grey area between the Data, Application and Infrastructure domains, particularly when identifying the appropriate technologies for data connectors. Topics such as API or IoT interfaces may fall within this border zone. We will delve into these details further in our next discussion.

Similarly, just as water can be collected temporarily or permanently and consumed, even can be supplied without being collected, ingested data can be temporarily or permanently stored in repositories like data lakes or supplied directly for immediate use. Through robust Data Management practices, this data can seamlessly integrate into storage systems and business processes or be directed to a streaming analytics platform. To harness the potential of ingested data, a variety of user tools, processing engines, and performance optimisation techniques are indispensable. These components are grouped within a Data Lab, facilitating efficient data exploration, analysis, and experimentation.

Organisations often have responsibility to provision datasets not only internally but also externally to meet the diverse needs of stakeholders. The EDRA wouldn’t be complete without provision patterns making datasets available across ecosystems. File-based Access patterns, encompassing traditional methods like SFTP and modern approaches such as cloud storage services, facilitate data sharing and consumption. APIs and web services play a pivotal role in enabling programmatic access to data, offering flexibility and scalability. Data virtualisation techniques further streamline access to distributed datasets without physical data movement, enhancing agility and efficiency. Additionally, real-time data provisioning involves publishing data feeds that continuously deliver updates to applications and analytics platforms, ensuring timely decision-making based on current data.

Much like consumable products made with water, such as bottled mineral water, data products contain specific information tailored to meet the needs of individual consumers. These can encompass various forms, including curated learning datasets for Artificial Intelligence (AI) models, customised dashboards for customers to track key performance indicators, targeted marketing insights for decision-makers, or even publicly published company reports.
Data sharing platforms, also known as Marketplaces, are relevant to both, Provision Patterns and Presentation systems within a reference architecture. These platforms facilitate secure and compliant data sharing with external parties. While these platforms are often owned and operated by the organisation, they can also be externally provided as a service. For example, Dataplace is a newly established government platform that allows users to request Australian Government data. Marketplaces are strategically positioned within the presentation layer of our EDRA diagram, facilitating seamless access to information.
All capabilities within the presentation layer are supported by well-designed processes and applications, ensuring timely and secure access to information, internally and externally.

Finally, this leads us to Level 2 diagram depicted below.

EDRA is to be tailored to the unique objectives of each organisation, evolving alongside the strategic planning efforts to accommodate additional details. In our next time discussion, we will delve deeper into the level-2 elements to provide a more comprehensive understanding of these capabilities.
Establishing a glossary to describe each element and assign statuses can enhance communication and collaboration among stakeholders. By mapping EDRA elements to both current and future state technology, organisations can leverage the architecture as a powerful tool for gap analysis and prioritisation.
Let’s clarify using an example. The No-SQL element within EDRA can be described in the Glossary as:
– ‘NoSQL engine is a non-relational database engine for storage and retrieval of data that is modelled in means other than the tabular relations used in relational databases. The data structures used by NoSQL databases are key-value, wide column, graph, or document, which make Inset/Delete operations faster in NoSQL.’
By overlaying current and strategic technology mappings, an organisation can identify the need to transition from one type of NoSQL technology to another. This transition will require planning and prioritisation among other identified gaps.
Copyright © NeatenMyData

