Home / Project Guides Internet technology architecture

High Availability Disaster Recovery

High-availability is not “server-free”, but core operations can continue or resume within the agreed time frame when hardware, network, applications or human operations are abnormal.

How does the enterprise core system work? Disaster tolerance, backup and failure exercises

First define the business acceptable interruption and data loss

Enterprises should define the objectives of availability, recovery time target RTO and recovery point target. Payments, transactions and internal query systems require different inputs.

Without business classification, there is often over-building of non-core systems, while genuinely critical chain protection is inadequate.

Remove single points from the entrance to the data layer

The load balance, application multiple examples, cache clusters, news clusters and database hosts constitute the basic high-availability links. Deployment should also take into account the impact of machine rooms, available areas and network failure.

Redundancy is not equal to availability. Fault switching, data consistency and reliance on service time overtime require clear design.

The backup must be restored and the disaster must be reversible.

Backup policy should cover databases, files, configurations and key keys, and set up off-site copies, retention cycles and access rights. More important is to resume validation on a regular basis to confirm that the backup is not " looks successful ".

Core systems can be built to both live and be isolated from the city, but the highest specifications should not be pursued blindly, based on the choice of business values and recovery targets.

Turning the program into a capability through surveillance and exercises

The monitoring should cover user experience, operational indicators, applications, infrastructure and external dependence.

Periodic break-nets, nodal failure, database switching and backup recovery exercises are conducted to detect gaps between the document and the real environment.

  • Documenting problems identified in the exercise and improving the responsible
  • Reverse average detection and recovery time
  • Update emergency response plans and contacts on an ongoing basis
Implementation table

Moving from reading conclusions to project input

The most likely problem after reading methodological articles is the acceptance of principles, which are not translated into the next step. It is proposed that the head of operations organize a 60-90-minute mini-workshop, choosing only one real process and not rushing to discuss the full platform.

Step 1: Establishment of a current status and sample baseline

The data are not used to set a good rate of savings, but to reverse the data.

Step 2: Clarifying the initial closure and inaction

The first phase is designed to allow a chain to run and be retraceable, rather than to build disaster backup, failure exercise, system stability into the same version.

Step 3: Match technical results to engineering evidence

The structure determines the need for a tracking relationship between the number of the demand, sample number, test results and version around “backup must be recoverable, and the disaster must be transposable”. The structure is based on the validation of capacity, peaks, availability, recovery time, frequency of distribution and failure data, avoiding the premature introduction of complexity beyond the team capacity for technologically advanced purposes.

Step 4: Receiving, inspection and disking with the same calibre

Assuming that the original process handles 600 tasks per month, an average of 20 minutes and a return rate of 10 per cent, the target can be stated as “six weeks after the start of the line, with a reduction of 25 per cent on average, and a return rate of no higher than the original baseline, given the relative complexity of the task.” This set only demonstrates the measurement method and does not represent any client outcome; formal indicators must be identified by the enterprise on the basis of its own sample.

  • Operational material: flowchart, role, sample mission, current issues and baseline data
  • Technical material: system inventory, interface, data access, deployment environment and security requirements
  • Project material: first-phase scope, exclusions, liability matrix, milestones and change mechanisms
  • Receiving and inspection material: test set, execution records, list of deficiencies, indicator queries and handover documents

When these materials are identified jointly by both the operational and technical parties, the method in the article is actually entered into the project. If key data, interface authorization or the responsible person are not in place, the logical next step is usually a limited diagnostic or PoC, rather than an immediate commitment to complete the work period and fixed total price.

Core elements

Implement methodology to project action

  • Decision input with RRO, RPO and business hierarchy
  • Redundancy, backup, disaster management and surveillance are essential.
  • Unrehearsed recovery programmes cannot be considered effective.
Related issues

Continuing to reconcile common issues in project decision-making

Business Info, Systems integration and Transport

How do third party API integrated and multi-system interface development generally offer?

The interface project cannot simply be quoted by the number of interfaces, as the same interface may be simply a query, but may also assume transaction, retest, reconciliation and security responsibility. The cost depends on the quality of the document, the test environment, field conversion, synchronization frequency, unusual compensation, performance and online support. It is recommended that the number of URLs be assessed by business links rather than counting only. The unknown interface can be technically validated and then formally quoted.

View full answer
Corporate information selection, integration and data governance

Can the API interface be fully compatible without a file?

Sometimes, but costs, risks and time increase significantly, and no certain connection can be promised. Teams need to confirm whether there is a legal mandate, test environment, logs, sample requests and original support.

View full answer
Corporate information selection, integration and data governance

How do you monitor interface failure and data discrepancies after systems integration?

The interface returns successfully and does not amount to a business process completion, and systems integration must monitor both the technical state and the results of the operation. Each request must have a unique tracking number, recording the source, target, state, time-consuming, retry, and business unit number. Payments, orders, inventory, etc., are also regularly reconciled. Aberrants must be entered into a retried, reimbursable or manual processing queue and not remain in the log.

View full answer
Contracts, payments, changes and project delivery

What information is required for the software project acceptance and inspection?

The objective of the information is to demonstrate that the system meets agreed standards and that the client can continue to operate and take over.

View full answer
Professional services for ZhiHua Tech

Need for further analysis in the context of the current state of the enterprise?

We provide IT technical advice, enterprise information construction, Software Project Outlook, product design, R & D delivery and systems delivery services.

Liaison consultants
Content liability statement

The publication body: Shanghai, like the ZhiHua Tech. This paper is used for technical and project decision-making purposes; facts, data and external perspectives are presented on page and can be verified in scope and do not constitute a commitment to the results of a specific project.Checking content clearance, source of information and correction policy

Extending Reading

More Internet technical architecture articles

Enter the topic 's front page