Private AI / buying guide
How to size an on-premises AI system
Start with the work you want the machine to do. The right configuration depends on the selected model, the amount of material in each request, and how many requests run together. Employee count alone cannot determine the hardware.
Describe the inputs and outputs
List your most frequent tasks and one demanding example. For each, record the input format, typical size, expected output, and acceptable waiting time. A short email draft and a long document comparison are different workloads.
Specify the capabilities you need: text generation, document search, image and scan interpretation, transcription, or media generation. Each capability needs an appropriate model and a supported workflow. A language model’s presence does not establish that every image, audio, or video task is included.
- Which file types and programs do staff use?
- How long are typical documents, recordings, or batches?
- Which outputs need a staff member’s approval?
- Which tasks must complete interactively, and which can run in a queue?
Memory, model size, and working space
A model needs memory for its weights as well as space to process requests. Document length, conversation history, and concurrent requests affect the working memory needed. Storage capacity and model memory serve different purposes: keeping a large archive on disk does not mean the full archive fits into a single prompt.
Quantization stores model weights at lower precision to reduce memory requirements. The format, runtime, and hardware must support the chosen method, and output quality still needs evaluation. A memory-fit estimate is not a throughput benchmark.
SPLAY’s catalog lists installed memory for each reference configuration. The technical appendix provides estimates for published open models. Consultation confirms the selected SPLAY model and workload on the proposed hardware.
Total staff and simultaneous use
Record both the number of people who need access and the busiest expected period. Twenty staff reading and editing documents do not necessarily create twenty simultaneous generation requests. An automated batch can create sustained load even when only one person starts it.
Use a representative trial to check response quality and waiting times while multiple requests run. Record the exact hardware, model version, settings, and test inputs. Results for one model should not be carried over to another model without testing.
What to bring to the consultation
Prepare a budget range, your preferred installation timing, and a description of your network and available equipment space. Name the person or IT provider who will operate the system after handoff.
List the software you use, such as Word, Excel, Photoshop, a document management system, or a case management system. Identify whether the work can use file exports or requires a specific connection. Tool access and integrations must be agreed separately.
SPLAY uses these requirements to select the hardware and model and define installation, training, and optional support. All configurations are pre-order only and by consultation only.
- A short description of the most important tasks
- Typical file sizes and formats
- Total staff and expected simultaneous requests
- Required model capabilities and software connections
- Budget range, timing, and support needs