Skip to main content
GJ
← All projects
Feb 2026 – PresentNew York, NY

Self-Hosted NAS, Virtualization & AI Agentic Infrastructure

A home lab engineered like production: redundant storage, containerized services, segmented remote access, and GPU-accelerated local LLM inference driving automated research and reporting.

  • Docker
  • Linux
  • ZFS/Btrfs-style redundant storage
  • GPU inference
  • Local LLMs
  • Python
  • Automation frameworks
  • Network segmentation

Architecture

Personal infrastructure built to production standards — the same architecture, security, and reliability discipline I apply at work, applied to a system I own end-to-end.

Architecture — abstract

  1. L6Hardware

    Compute, GPU, redundant storage

    • I/O· Storage
    • compute· Containers
  2. L5Storage

    Redundant pools + automated backup

    • I/O· Hardware
    • reports· Automation
    • volumes· Containers
  3. L4Containers

    Docker isolation for every service

    • compute· Hardware
    • volumes· Storage
    • segmented access· Networking
    • runtime· AI
  4. L3Networking

    Segmented zones, minimal exposure

    • segmented access· Containers
  5. L2AI

    GPU-accelerated local LLM inference

    • runtime· Containers
    • inference· Automation
  6. L1Automation

    Agents for research & reporting

    • inference· AI
    • reports· Storage

Layers are abstract roles, not specific hosts.

Flows show logical data classes, not addresses.

Problem

Commercial cloud AI services put usage caps on exactly the workloads that matter, cost scales with every experiment, and sending research data to third parties is a non-starter. The system needed to run heavy local LLM inference and autonomous agent workloads continuously — while still being a safe, fault-tolerant place for personal data.

Approach

Treat the lab as a layered system with explicit boundaries. Storage is redundant and independently backed up so no single failure — disk, container, or operator — loses data. Services are isolated in Docker containers so dependencies can't bleed between workloads. Remote access is segmented: exposed surfaces are reduced to the minimum, with internal zones walled off from anything reachable from outside. GPU inference is benchmarked across models and configurations, then tuned for latency and throughput rather than peak specs. On top of that, AI agents are wired into the automation layer so recurring research and reporting jobs run themselves and hand back structured output.

Results

  • Redundant storage with automated backups — no single point of failure for personal data
  • Containerized service platform with isolated dependencies and repeatable deployment
  • Segmented remote access with a minimal external footprint
  • GPU-accelerated local LLM inference tuned across benchmarked models for latency and throughput
  • Autonomous AI agent pipelines for research and reporting
  • Roughly 100+ hours per month of manual work reclaimed by automation

Lessons learned

  • Benchmark with your own workloads — vendor and synthetic numbers diverge from real inference latency.
  • Segmentation pays for itself: the effort spent walling off zones is cheaper than any incident.
  • Backups are a system, not a setting — restore drills are what make them real.
  • Agents need guardrails and structured outputs; free-form automation drifts.