Editor’s Note:
In the age of digital transformation, the commodification of data elements, and the pursuit of high-quality data, the value of data has surged. Whether it’s financial institutions or fintech companies, each has its unique interpretations and practices regarding data, from the fundamental data architecture to advanced data applications. Sunline Technology, consistently attuned to the nuances of big data, initiates this discussion on “data architecture” to kickstart a series of articles. Together, let’s explore data perspectives that resonate with the contemporary industry. Stay tuned.
Author | Sunline Technology Big Data Research Institute
Content | This article comprises 3170 words and is estimated to take 10 minutes to read.
For quite some time, the concept of “data architecture” remained a niche area of expertise. In China, efforts related to data governance were primarily centered on establishing data standards.
However, a pivotal moment occurred on February 9, 2021, when the People’s Bank of China introduced the “Guidelines for Building Data Capability in the Financial Industry.” This marked a watershed moment, as it officially expanded the understanding of data architecture to encompass the entire industry.
This transformation entailed a shift from “data standards” to “data specifications,” symbolizing a profound change in perspective.
What is Data Architecture?
As per ISO/IEC/IEEE 42010:2011, architecture refers to a system’s fundamental organization, seen in its components, relationships, environment, and guiding design principles.
“Architecture” originally comes from construction. While ancient cottages didn’t need detailed designs, modern skyscrapers rely on blueprints.
So do Data management, it needs architecture. It’s the blueprint for managing data assets, including models, definitions, mapping specs, flows, and structured data interfaces (DAMA DMBOK2). Data architecture needs ongoing maintenance, like physical architectural design.
GB/T36073-2018 “DCMM Data Management Capability Maturity Assessment Model” describes data architecture. It defines data requirements, guides data asset control and integration, deploys data environments, and manages metadata.
How do we interpret data architecture, its components, relationships, and associated environments?
Within enterprise data architecture design, key components encompass the creation of an enterprise data model and the design of data flows (as outlined in DAMA DMBOK2). These are foundational elements of data architecture.
Constructing an enterprise data architecture from the ground up demands substantial time, financial resources, and entails inherent risks. Fortunately, the realm of practical enterprise data architecture benefits from the existence of well-established industry data models that serve as valuable references. Consequently, in most scenarios, the most comprehensive data architecture design document assumes the form of a formal enterprise data model. While the physical data model certainly constitutes a data architecture document, it’s worth noting that it arises as a product of data modeling and design rather than falling directly under the purview of data architecture.
Incorporated within the data architecture capability domain of DCMM are several critical facets, including data models, data distribution, metadata management, data integration, and data sharing capabilities. The standards outlined in DCMM have garnered extensive recognition within the industry, with some prominent organizations achieving level four or five certification in data management capabilities.
At the heart of data architecture lies the data model, shaping data distribution and integration. Data distribution feeds into data integration, and, conversely, data integration imposes requirements on data distribution for efficient interaction.
Metadata plays a pivotal role as well. Enterprise data model design, data distribution, and data integration, all encapsulated within metadata, illuminate the intricate relationships between these components. Without this metadata, the intricate web of connections remains concealed, rendering the data inaccessible to users.
Data architecture exists as an interconnected element within the broader enterprise architecture landscape. Examining this interplay, we find that data architecture shares a close affinity with enterprise application architecture and technology architecture.
Application architecture serves as a foundational input to data architecture. The distribution of data in business applications and data application systems hinges on the functions these systems perform. Application architecture essentially dictates data distribution, thereby influencing critical aspects of data architecture such as data definition, integration, and interaction. Yet, data architecture should not merely conform to the inputs of application architecture. Instead, the data model design within the application system should align with enterprise data architecture, ensuring data distribution and integration principles influence the functional distribution of application systems.
Conversely, data architecture feeds into technology architecture. Requirements concerning data integration, interaction, and storage planning find their articulation in technology architecture. Technology architecture, in turn, acts as the foundational bedrock supporting data architecture, with decisions regarding database software, data storage, and other considerations potentially shaping application architecture and data architecture choices.
Three decades since the inception of the China Financial Standardization Committee, remarkable progress has been made in standardizing financial data, positioning the financial industry at the forefront of data management. However, a closer examination of practical experiences in data management reveals a set of pertinent issues.
1. Oversimplified Data Architecture Management:
In the age of affordable storage, some companies have adopted an indiscriminate approach to data collection, emphasizing quantity over quality. This undiscriminating “large and all-inclusive” data collection mindset neglects careful data differentiation, Total Cost of Ownership (TCO), and Return on Investment (ROI), resulting in unchecked data expansion.
2. Misalignment Between Data Management and Enterprise Data Architecture Governance:
Several enterprise data architecture initiatives primarily focus on data analysis platforms (data warehouses or data hubs) and data integration, lacking comprehensive data models. This oversight overlooks input from enterprise data distribution (application architecture), potentially leading to the use of non-authoritative data sources and data quality issues.
3. Incomplete Data Governance Development:
While some companies have established data standards, their implementation remains lacking. Recognizing and addressing issues like “synonym different name, same name different meaning” and standardizing terminology and concepts have not been integrated into data architecture, hampering substantive progress.
4. Overemphasis on Availability Over Quality:
In the quest for higher availability, some financial enterprises, driven by internet solutions, have raised tolerance for data inconsistencies and quality concerns. However, the financial industry’s stringent requirements for financial statements and regulatory compliance demand far higher data quality standards than the internet sector.
Path Forward: Embracing an Enterprise-Level Data Management
Effective data management calls for an enterprise-level architectural perspective. Often, data is possessed without a corresponding data architecture. Data architecture typically takes a backseat, only gaining attention once technology platforms are chosen. Application systems frequently overlook the fundamental requirements of data architecture. When technical platforms struggle to meet customer expectations and fail to provide a satisfying experience, a reevaluation becomes essential, involving both data architecture and application architecture.
1. Embracing a Top-Down Perspective:
Enterprise data architecture, spanning data governance, data asset management, data warehouses, data hubs, and digital transformation, necessitates adopting a top-down perspective at the enterprise level and conducting systematic architectural design. At each stage of the data lifecycle, from design to usage, stakeholders must assume responsibility for the data products they engage with, considering application and technology architecture on both sides of the spectrum.
2. Staying Rooted in Abstract Thinking:
Achieving a stable data architecture involves abstract thinking that transcends the intricacies of data. Resolving issues like “synonym different name, same name different meaning” is just the initial step in the extensive journey of data architecture management. Comprehending the essence of business, abstracted from the noise of data, and expressing it in a structured and systematic manner forms the basis of a robust framework.
3. Fostering Collaboration with Data Governance:
Data governance stands as another crucial component influencing data architecture. The ultimate objectives of data architecture can only be realized through harmonization with data governance. While data architecture ensures the correct execution of tasks, data governance ensures that the right tasks are performed, in accordance with data architecture.
4. Leveraging Enterprise Data Models:
Starting with the inception of an enterprise conceptual data model, stakeholders must recognize the pivotal role of enterprise data models. Systematic data governance must seamlessly integrate with data modeling, refining it into an enterprise data model that guides data distribution and integration across various enterprise systems. This supports data reuse and agile development within application systems.
5. Managing the Entire Data Asset Lifecycle:
Fostering a data asset perspective treats data as a valuable asset. It begins with the inception of data, cultivating cost and benefit awareness throughout its lifecycle. All stakeholders, from data designers, producers, processors to managers, should understand that they invest in data throughout its journey. Optimizing the data management ecosystem, attaining enterprise-level data integration, and maximizing data asset value become paramount
6. Establishing a Robust Data Analysis Ecosystem:
Structured and unstructured data possess varying value densities, necessitating distinct processing techniques and workflows. Recognizing these variations, a comprehensive data analysis ecosystem should be constructed. It can effectively manage diverse data forms and lifecycle stages, ensuring high-quality data and a user-friendly experience across the board.
In Conclusion
The management of enterprise data is an ongoing, evolving endeavor, deeply intertwined with an enterprise’s growth and transformation. As businesses adapt and change, data architecture must progress in lockstep. The mere implementation of a data hub does not guarantee enhanced data management capabilities or the delivery of high-quality data to users. To address these challenges, it’s essential to adopt an enterprise-level architectural perspective, establishing robust data governance, and embracing abstract thinking to build a stable and efficient data architecture foundation.