The DIY Guide to Secure, Local AI

A step-by-step blueprint for building an offline, air-gapped AI workstation.


// 01. REQUIREMENTS

What Do You Need? (Defining Your AI Workload)

Before selecting hardware, translate your business tasks into technical requirements.

Healthcare & Administration:

Summarizing patient notes, extracting data from intake forms, or processing handwritten records requires a fast, Vision-capable model.

A lightweight multimodal model (7B–11B parameters) handles real-time OCR on 1–2 page documents without the heavy memory overhead of larger systems.

Legal & Compliance:

Cross-referencing 50-page contracts, analyzing case law, or performing document discovery requires a massive model with a large context window (the ability to hold many pages in memory simultaneously).

Models in the 32B–70B+ parameter range with 128k+ token context windows retain full document memory at once, preventing dropped clauses and hallucinated citations.

Engineering & R&D:

Querying proprietary CAD documentation or technical manuals requires deep reasoning capabilities and high-precision local RAG (Retrieval-Augmented Generation).

A mid-sized reasoning model (14B–32B parameters) paired with a local vector database delivers precise, source-cited answers from complex technical data without exposing your IP to the cloud.

How to Calculate Your Office’s AI Throughput Needs

To select the right hardware, estimate how many discrete AI tasks (entries) your office needs to process per hour. An "entry" is one complete cycle: uploading a document, processing it, and generating the final output.

1–5 Entries per Hour: Compact Integrated Workstations (Small Form Factor)

Best for independent professionals or light administrative use. Ideal for intermittent, single-user tasks where waiting a few extra seconds for generation is acceptable. These entry-level local nodes handle low-concurrency workloads efficiently using shared system memory, without the need for a dedicated, high-power GPU.

5–50 Entries per Hour: AMD Radeon RX 9060 XT

The standard for small teams or steady, moderate daily workflows. This tier provides dedicated AI acceleration, ensuring fast token generation and smooth handling of concurrent requests from 2–3 staff members without system lag.

50–250 Entries per Hour: AMD Radeon RX 9070 XT

Designed for busy departments or high-volume document processing. This high-performance tier handles heavy batch processing, complex reasoning chains, and simultaneous multi-user requests with minimal latency, keeping fast-paced teams productive.

250+ Entries per Hour: AMD Radeon AI PRO R9700

The enterprise-grade solution for mission-critical, continuous operations. Built for sustained, heavy computational loads, this workstation tier guarantees maximum throughput, advanced driver stability, and zero thermal throttling during all-day, multi-user, high-concurrency environments.

Your workload determines the required model size. Your model size directly dictates your hardware specifications: the specific GPU model, its VRAM (Video RAM), or the unified System RAM for Small Form Factor (SFF) Local Nodes. Finally, your peak hourly entry volume determines the overall hardware tier. Underestimating volume leads to workflow bottlenecks, while overestimating wastes budget. We help you calculate your exact peak load to ensure your system scales perfectly with your business.


// 02. The Brain

Choosing the AI Model

Selecting the right local AI model requires balancing two factors: Capacity (how much data it can hold in memory) and Compliance (how securely it handles that data under GDPR and professional guidelines).

Step 1: Choose Your Capacity (Model Size & RAM/VRAM)

Up to 16GB (Lightweight | ~7B–14B Parameters):

Best for fast, single-task operations. Ideal for summarizing 1–2 page documents, basic intake forms, or simple administrative queries. Low latency, high speed.

Up to 32GB (Mid-Weight | ~14B–32B Parameters):

The standard for professional workflows. Required for multi-document reasoning, standard local RAG (Retrieval-Augmented Generation), and analyzing 30–50 page contracts or technical manuals without losing context.

64GB+ (Heavy-Weight | ~32B–70B+ Parameters):

Built for complex, multi-step logical reasoning. Necessary for massive context windows, cross-referencing hundreds of pages of case law, or querying extensive proprietary CAD and engineering databases simultaneously.

Flexible Model Selection

Because the open-source AI landscape changes rapidly, your system comes preloaded with a selection of open-source (MIT/Apache 2.0) models via Ollama. This allows you to test and choose the model that best fits your specific workflow. Our baseline selections are guided by independent, community-tracked benchmarks—such as the Onyx Self-Hosted LLM Leaderboard—ensuring every preloaded option is proven to run efficiently on local, air-gapped hardware without cloud dependencies.

Step 2: Choose Your Security Posture

Aligned with guidelines from the Law Society of Ireland, HSE data governance standards, and Chartered Accountants Ireland, every AI deployment must be classified by its level of data and physical control:

  • Private AI Deployment
  • Your AI runs in an environment you exclusively control, ensuring your data is never exposed to public AI providers. This deployment can run on a rented virtual server, a private cloud, or physical hardware you own on your premises. The defining feature is strict data control, not physical location. This satisfies GDPR and professional confidentiality requirements for solicitors and accountants, ensuring your inputs are never used for third-party model training.

  • On-Premises AI
  • The AI runs on physical hardware you own, located either directly in your office or in a dedicated data centre rack. This provides absolute physical possession of the infrastructure. However, physical ownership alone does not guarantee security; it must be properly configured as a "Private" deployment to ensure network isolation. The Law Society of Ireland specifically recommends this level of physical control to protect legal privilege and client confidentiality.

  • Air-Gapped AI
  • The highest level of security. The on-premises hardware is physically and logically disconnected from all external networks, including the internet and corporate intranets. Data enters and exits only through strictly controlled physical channels. This eliminates all network-based risks and is the required standard for handling highly sensitive special-category data, such as comprehensive patient medical records or critical intellectual property.

The Rule of Combination:

Capacity and security are not interchangeable. A 70B-parameter model is useless if it is Red (cloud-based), and a Green (air-gapped) system will bottleneck your firm if it only has 16GB of RAM for a heavy legal workload. We configure the exact intersection of your required capacity and your mandatory security tier.


// 03. The Body

Sourcing the Hardware

Lorem ipsum dolor sit amet consectetur, adipisicing elit. Debitis omnis deleniti rem, numquam quia quibusdam itaque exercitationem vitae voluptas vel accusamus perspiciatis rerum totam neque, nam doloremque a, delectus quas.

// 04. The Nervous System

The Software Stack

Lorem ipsum dolor sit amet consectetur adipisicing elit. Odit optio ipsam dolorum quaerat laudantium repellat, aliquid amet doloribus, corrupti reiciendis cupiditate totam eaque et obcaecati libero harum! Quasi, officiis repudiandae?

// 05. The Proof

Security & The Audit Trail

Lorem ipsum dolor sit amet consectetur adipisicing elit. Fugit eum, fugiat cupiditate odit quo explicabo consectetur non? Voluptatem, ducimus. Aperiam, nulla voluptatibus. Rerum, at ad! Quia optio sint veritatis odio!