boundary.labs
Boundary Labs  /  Research Infrastructure

The Stack

Hardware specs, inference configuration, optimization history, and agent infrastructure. Everything documented. Raw numbers only.

New field report: older laptops, Omarchy, and the daily-work improvements — September 6, 2026.

Compact local lab: an always-on orchestration node carrying the agents and services, plus commodity laptops for development and CPU-inference research. The dedicated GPU inference node below was retired 2026-07-20 (GPUs sold) — its specs stand as the documented tower-era configuration; production inference now runs through the same internal gateway against an API backend.

GPU Inference Node — tower era, retired 2026-07-20

SystemCyberPowerPC GXi3400BSTV17
CPUIntel Core Ultra 7 265F
CPU config20c/20t  ·  5.3GHz boost
GPU ×2RTX 5060 Ti 16GB GDDR7
ArchitectureBlackwell SM_120
VRAM total32 GB GDDR7 (TP=2)
PCIe configx8 Gen5 + x4 Gen4
RAM32 GB DDR5 (192 GB max)
Storage2 TB PCIe 4 NVMe
CUDA13.0.3
InferencevLLM + llama.cpp
GatewaySingle internal inference gateway
OSUbuntu 24.04

Orchestration Node

SystemBeelink EQI12
CPUIntel i5-1235U (12th Gen)
CPU config12c  ·  4.4GHz boost
RAM64 GB DDR4 (upgraded 2026-08-27)
Storage1 TB NVMe
OSopenSUSE Tumbleweed (since 2026-07-08)
UptimeAlways-on
Active servicesharness, Mike, Hermes, autoresearch, chronicle, localfamouscoffee
Webnginx  ·  Cloudflare tunnel
NetworkWired LAN
Remote accessPrivate mesh access
Backup476 GB NVMe  ·  3am daily  ·  30-day retention

Dev Workstation  — CPU inference testbed

SystemDell XPS 15 9570
CPUIntel i7-8750H (6c/12t, no AVX-VNNI)
RAM32 GB DDR4
RolePrimary dev  ·  CPU agent-fitness eval host
Remote accessPrivate mesh access

Network  — Topology

Cluster linkDirect Ethernet
ExternalPrivate mesh access
Inference accessSingle internal gateway for all clients
Web routingCloudflare tunnel → nginx
SSHPrivate-only, not exposed publicly

Production backend: GPT-5.6 Luna (OpenAI API)
Verified 2026-09-06. Agent orchestration and memory run locally; production inference uses an external API through a shared gateway. This is a dated verification, not a live status feed.

The configurations below are historical GPU results from before the tower retirement on July 20, 2026.

93
tok/s — dated API measurement
2026-08-30, median of 5 × 600-token runs via :8010, reasoning disabled; observed range 92–102 tok/s
2,436
tok/s prompt (dual GPU)
tower era — dated lab peak
Retired
Dedicated GPU inference tower
since tower retirement 2026-07-20
5.7
tok/s CPU agent eval (30B MoE)
Dell i7-8750H — 2026-08-24, details
Model Quant Gen Context Server Status Notes
Qwen3.6-27B GPTQ INT4 + MTP n=3 (Genesis) AutoRound INT4 ~97 t/s warm (2026-07-08) 65K vLLM (Genesis-patched) production until tower retirement 2026-07-20 Final local production backend — restored 2026-07-03 after the Ornith trial. GPTQ-Marlin, TP=2, MTP n=3 speculative decoding, chunked prefill + prefix caching.
Ornith-1.0-35B Q4_K_M GGUF / AEON NVFP4 ~101 t/s warm / 124 peak (prod config) 131K (NVFP4) llama.cpp / vLLM 0.23 trial ended 2026-07-03 Production trial 2026-06-26 → 2026-07-03. Lab peak 129.9 t/s short-context (2026-06-25). Won sequential tool composition vs DeepSeek V3.2 (20/20 vs 11/20, ~10× faster); DeepSeek won the suite overall by one point.
Nemotron 3 Nano 30B (MoE hybrid) Q4_K_M 117.34 t/s med / 117.60 peak 32K llama.cpp cuda128-clean historical — measured 2026-05-20 Mamba/SSM hybrid MoE. ~3B active params/token. Requires cuda128-clean build; cuda13 crashes with MMQ CUDA error on SM_120.
AEON (Qwen3 NVFP4) NVFP4 (ModelOpt) ~69 t/s 122K vLLM 0.9+ stopped Blackwell-native quantization. MTP speculative decoding n=3.
Qwen3.6-27B (SSM hybrid) Q4_K_M ~22 t/s 65K llama.cpp stopped Mamba/SSM hybrid. Fast prefill (960 t/s), slow gen (SSM bottleneck).
Gemma 4 26B (CPU baseline) Q4_K_M ~11.7 t/s 32K llama.cpp historical CPU-only orchestration-tier baseline from the pre-GPU phase. Consistent swap usage 4–8 GB.
Blackwell SM_120 fast CUDA kernel paths: q4_0 and f16 KV only. q8_0, q5_0, iq4_nl all degrade significantly. NVFP4 (ModelOpt) is Blackwell-native quantization — 4× smaller KV footprint vs bf16, requires CUDA 12.8+. This system runs CUDA 13.0.3; Qwen3-series NVFP4 (FlashInfer CUTLASS SM_120 FP4 kernels) should now be unblocked — previously blocked on CUDA 12.8. llama.cpp users: build target matters. cuda13 crashes with CUDA MMQ error on SM_120; cuda128-clean required for llama.cpp GGUF inference.

130+ automated autoresearch runs across six model architectures using a scripted loop — a subset of the 215+-run total benchmark record. Stopping criterion: +5 tok/s improvement per iteration. Full logs at inference-research.

Key Configurations — Qwen3.6-35B-A3B MoE (dual GPU, April 2026 campaign)

ConfigurationGen SpeedPrompt SpeedDelta vs BaselineNotes
Dual GPU, all-on-GPU, f16 KV 107 t/s 2,436 t/s +50% vs single GPU Final config. TP=2, PCIe x8+x4.
Single GPU, all-on-GPU, f16 KV 71 t/s — Pre-dual baseline After CPU offload fix. Clean VRAM fit.
Full GPU offload (Exp 2 breakthrough) 70 t/s 222 t/s +118% gen, +200% prompt vs expert-on-CPU config.
Expert tensors on CPU (MoE routing) 32 t/s 74 t/s −55% CPU bottleneck on MoE expert routing.

CPU Inference Baseline — Orchestration Tier (17-day log)

MetricValue
Sample periodApril 2–18, 2026
ModelGemma 4 26B Q4_K_M
Avg gen throughput10.49 tok/s
Range5.02 – 11.72 tok/s
Avg time to first token616 – 2,224 ms
Avg swap used4.6 – 7.7 GiB
Speedup vs GPU tier6.6× (10.5 → 74 tok/s)
Swap usage is the operationally significant CPU-tier finding. In the earlier multi-agent phase, co-locating several always-on processes and local inference on a 32 GB machine consistently drove swap. The architectural fix — separate inference and orchestration nodes — eliminates this entirely. Not a tuning problem; an architectural one.

Autoresearch Results by Category

Optimization delta from experiment baseline (MoE, dual GPU)
full GPU offload (MoE experts)
+118%
dual GPU (TP=2)
+50%
f16 KV cache
+12%
q8_0 KV cache
−18%
expert tensors on CPU
−55%

Core agent processes run as hardened background services. Auto-restart, post-boot recovery, and gateway-based inference routing are all part of the runtime design. Slack Socket Mode remains the primary interactive path for the agent stack.

ProcessFrameworkInterfaceModel tierStatus
Agent Harness Python / custom Slack Socket Mode Internal inference gateway live
Mike RelayV3 (Python) Discord + Telegram + Slack + IRC Internal inference gateway live
Harness Personas Internal harness roles Not separately exposed Internal gateway / task-dependent internal
Hermes Python / hermes-gateway.service Gateway agent — inference via internal :8010 proxy Gateway (GPT-5.6 Luna) live
Chronicle Python / cron 2am daily cron Claude API live

Service Hardening Details

Crash recovery and boot recovery are actively hardened, but the exact unit settings evolve with the live system. The durable design claim is local-first orchestration: the agent runtime, memory, and gateway stay on local infrastructure, while the inference backend behind the gateway is an operational variable — an API backend; see the verified status above, previously local GPUs, swappable without touching any client.

All benchmark data and logs referenced in the research paper are available publicly.

ArtifactLocationContents
Live benchmark data boundarylabs.org/benchmarks.html Inference metrics, MESA scores, LongMemEval results
Raw experiment logs inference-research 215+ raw benchmark runs, drivers, and verdict docs per research program
Agent infrastructure randomchaos7800-hub Agent harness, benchmark tooling, selected repos
Research paper papers.html Full preprint text, tables, references
Weekly build logs dinovitale.com Weekly recap posts with build decisions and operational notes