First, give conclusions that can be used for decision-making
The side of the platform needs to consider whether Web and API services, work processes, databases, caches, object storage, vector retrieval, document resolution and log monitoring are separate. The side of the model depends on using cloud-end API, local small models or multi-model reasoning. The peaks are recorded at the same time, the size of the single upload file, the increasing volume of documents, the frequency of knowledge updates, the number of workflow nodes and the acceptable response time. The production environment also has to be set aside for backup recovery, disk growth, component upgrading and failure to switch space.
What conditions need to be identified before judgement is made?
The same question may have different answers under different business, data and project phases. It is suggested that the following conditions be checked and that the common findings on the web be incorporated into their own projects.
Suggested order of advance
First, we'll be clear about the target and the border.
(b) Collating the baseline of users, tasks, documents, interfaces and response time.
Validation Key Dependence
The platform, the knowledge processing and the model reasoning will be used separately.
Development of assessable outcomes
The pressure and capacity tests are performed using real files and workflows.
Make sure you decide the next step with the real results.
Production specifications based on P95 delays, queues, resources and failure results.
How do you understand it in the actual business?
An enterprise with 30 internal users, but with a large volume of PDF imported daily and multi-step workflows running, the knowledge analysis and back-office tasks may be more resource-intensive than hundreds of users who only occasionally ask questions and answers. If local model reasoning is requested, separate tests should be conducted on different models, context lengths, and the co-disposal and insulation, and a demonstration server cannot be considered a production configuration directly.
The easiest pit to step on.
Only estimated by number of registered users, without actual tasks and document loads
Successful deployment of Diffy equals production capacity achievement
No growth of queues, databases, disks and vector indices monitored
How should we end up receiving and confirming?
The acceptance and inspection shall be performed in the target environment with agreed and asked questions and answers, document uploads, updates of knowledge and workflow tasks, recording P50, P95 responses, queue waiting, error rate, CPU, memory, disk and model resources; and complete backup restoration, disk alarm, service restart and version back exercise.
When preparing to communicate with suppliers or internal teams, it is recommended that current processes, representative samples, existing systems, planning time and budget levels be brought. First, the unknown items are clearly marked, and then the decision is made to use diagnostics, PoC, fixed-range projects or ongoing research and development, which is usually more reliable than a direct demand for a price and duration without borders.