Background: The "Fat of Business Data"
Over the past decade, China has made great strides in the development of information technology for its enterprises - the ERP account, CRM, WMS, WMS, OA, and the big and small business systems cover one business area each. But one ironic reality is:The more systems there are, the more data there are, the more vague the global view management can see.。
Typical symptoms: Revenue figures exported from ERP by the finance sector are not matched by the amount of signed money from CRM statistics; stock turnover days for supply chain teams do not match the actual login and login data of warehouse WMS; CEO wants to see "the trend of Māori rates across product lines over the past 12 months" and IT teams need to coordinate three departments, run five statements, and manually consolidate before giving a "presumably possible" figure.
It's not a problem with data. It's a problem.Data scattered across the business systems "coastal", lacking harmonized definitions, standards and access points. The scenes described here are for businesses with multisystem data isolated, report calibrations inconsistently and analysis inefficient.ZhiHua TechData mid-station construction and data governance methodology do not represent the disclosure of data for specific clients.
Typical operational challenges
1. Data scattered multiple systems, global unified view missing
- ERP controls financial vouchers, CRM controls customer business, WMS controls the flow of water from stock - the data model of the three systems is independent of each other, with different codes, attributes and status definitions for the same "client", "product" and "order" in their respective systems.
- Management wants to see "Customer 360" — which products, transaction history, complaint records, returns — it needs to pull data manually from at least three systems and then collide them together, and the process takes time and results are prone to error.
- Each business maintains its Excel statements and database queries separately, and the same indicator (e.g., "monthly sales") may have five different figures on the computer of five people.
2. Data quality is uneven and analysis is low in credibility
- The material master data "one size multiple": The same screw of the same supplier, which is a set of codes in the ERP, another set in WMS, and the third set in the R & D BOM. This leads to frequent errors in inventory counts, procurement reconciliations and cost accounting.
- Repeated and conflicting client data: multiple records (slightly different names and different contact details) of the same customer in CRM for historical reasons, sales are not clear about which is the latest version and marketing may send the same text message three times to the same customer.
- Data missing and abnormal values: Key fields (e.g. contract amounts, delivery dates) are empty or clearly abnormal (date 1970-01-01), but there is no automated quality detection and repair mechanism, and dirty data flow to analytical reports.
Data analysis relies on manual and poor decision-making time-frame
- Business analyses rely on cousins' Excel puzzles overnight, which are three weeks ago. Competing parties have adjusted their pricing strategies to real-time data, and are waiting for results.
- When business asks an exploratory question – “What is the repurchase rate of new clients in the last three months?” – data analysts take half a day to write SQL, calibrate, and produce results. When they get the answer, business attention may have shifted.
- Lack of self-help analysis capacity: Operators cannot drag data for exploratory analysis themselves, and all data needs are queued in the IT sector ' s work list, creating typical data "bottleneck effect".
ZhiHua Tech Data Media Construction Methodological
Data asset inventory and master data governance
The first step in building the data center is not technical selection, but rather technical selection.Find out what the data are, where they are distributed, who's using it, and what quality they are.. ZhiHua Tech, through the data asset inventory workshop, works with business units in the enterprise to:
- Global Data Asset Directory: Combine data sheets, data fields, data update frequency and data for all business systems, and Owner, to form enterprise-level data maps. This step addresses the question of where the data are.
- Main data standardizationEstablishes enterprise-level coding rules, attribute specifications and maintenance processes around core data entities such as materials, customers, suppliers, organizations, accounts.
- Data quality baseline• Quantitative indicators (completeness, uniqueness, consistency, timeliness, accuracy) defining data quality, a full quality scan of existing data, identification of the most critical quality panels and development of governance plans.
2. Dataset formation and stratification modelling
The technical skeletons that build the data medium are set up after the data assets are clear:
- Dataset Layer (ODS): Synchronizes the original data from the business systems into the data medium in near real time through the CCDC (change data capture), ETL/ELT conduit and API. Keep the full raw data pattern, unprocessed, and ensure data traceability.
- Data Repository Level (DW): using the Kimball dimension modelling or DataVault methodology, which cleans, standardizes, links raw data and constructs analytically oriented thematic domain models - sales themes, procurement topics, inventory topics, financial topics, human subjects, etc. This layer is "a single source of credible data".
- Datasets (DM): For specific business scenarios (management cockpit, sales funnel analysis, supply chain performance analysis, customer image), construct a broad-sheet and indicator view of pre-aggregated, allowing business parties to obtain the required analysis at the minimum SQL complexity.
- Indicator calibration management: Create a dictionary of enterprise indicators – each indicator (e.g., GMV, Māori, stock turnover days, customer retention rates) has a single name, calibration, data source, and responsible person. When quoted in any statement, the index is taken from the dictionary of indicators, eliminating the problem of "unspecified numbers" at all.
3. Self-help analysis and data services
The value of the data center is ultimately reflected in "make data accessible to those who need it":
- Self-Help BI Platform: Operators connect directly to the dataset municipal level by drag-and-trucking BI tools (e.g. Metabase, Superset or Power BI), create self-generated reports and panels without IT intervention in writing SQL. IT team transitioned from "data picker" to "data platform carrier".
- Data API Service: External exposure of data assets in the middle channel via REST API for front-end applications, mobile end and small program consumption. For example, the CRM system can search for customer 360 image data in real time, rather than maintaining a copy of a customer that may expire locally.
- Data blood and impact analysis: Record the full data flow path (bloodline) from the source table to the end of each report. When a field of a table is changed upstream, all downstream dependent parties are automatically informed, avoiding a "upstream trip without knowing who is using it"
Key technical components
| Component | Annotations |
|---|---|
| Data set engine | A low-code data conduit based on Apache SeaTunnel / Flink CCDC, which supports batch synchronisation of 50+ data sources, including automatic Schema map and unusual data capture |
| Data Lake/Silo Storage | MinIO / HDFS as data lake base, StarRocks / Clickhouse as ORAP engine, PostgreSQL to manage metadata and indicator dictionaries, balancing storage costs with query performance |
| Data quality engine | Configured quality rule engine (non-empty, unique verification, domain validation, cross-table consistency ratio), time-scan and generate quality reports, and abnormal data automatically push to data Owner |
| Indicator management platform | Visualized indicator dictionary management, indicator blood tracking, indicator API automatic generation and indicator life cycle management (creation of a top-line change line) |
| Data security and access | Level-authority control, dynamic data desensitization (cell phone number, ID card, bank card number), operating audit logs to meet data security law compliance requirements |
Deliverables
| Phase | Delivery | Main elements |
|---|---|---|
| Data asset count | Data asset catalogue + quality assessment report | Global data sheet lists, field blood, data quality ratings, master data governance programmes and prioritization |
| Platform building | Data medium technology platform | Full deployment of data set-forming conduits, silo-strenching models, ORAP engines, BI tool integration and data API gateway |
| Data products | Management cockpit + self-help analysis platform | Core business indicator board, business theme data set city, self-help BI report template and indicator dictionary |
| Operating mechanisms | Data governance operational norms | Master data maintenance SOP, data quality monitoring rules, indicator management processes and training materials |
Intended value orientation
- One number, one caliber.The definition, source and calculation logic of the core enterprise performance indicator are managed centrally in the middle stage, and the figures that management sees on any report and on any screen are derived from the same data.
- From "T+7" to "T+0.": Business data update cycle is manually aggregated from week/month to hour or minutes to automatically synchronize, and decision makers can see and respond to data signals on the day of the problem.
- Release IT productivity: Operations accomplish 80% of their daily take-out needs by self-help BI, and the IT team focuses on building data platforms and improving data quality rather than being tired of taking a number of work orders.
- Set the data base for AI.: High-quality, well-managed business data are prerequisites for AI implementation. Once the data medium is completed, AI's RAG knowledge base build and model training can directly connect to clean data sources, significantly reducing the lead-up cycle for AI projects.
📎 Know more:
- Enterprise Data Platform solutions — one-stop data medium from data set-up to self-help analysis
- Information for enterprises - Integrated Information System Planning and Improgramming
- Free consultation — communication with ZhiHua Tech team on data governance needs
Need for further analysis in the context of the current state of the enterprise?
We provide IT technical advice, enterprise information construction, software project Outlook, FDE enterprise AI application and software product design and delivery services.