The Stack
Hardware specs, inference configuration, optimization history, and agent infrastructure. Everything documented. Raw numbers only.
New field report: older laptops, Omarchy, and the daily-work improvements — September 6, 2026.
Research Cluster
Compact local lab: an always-on orchestration node carrying the agents and services, plus commodity laptops for development and CPU-inference research. The dedicated GPU inference node below was retired 2026-07-20 (GPUs sold) — its specs stand as the documented tower-era configuration; production inference now runs through the same internal gateway against an API backend.
GPU Inference Node — tower era, retired 2026-07-20
Orchestration Node
Dev Workstation — CPU inference testbed
Network — Topology
Inference Configuration
Production backend: GPT-5.6 Luna (OpenAI API)
Verified 2026-09-06. Agent orchestration and memory run locally; production inference uses an external API through a shared gateway. This is a dated verification, not a live status feed.
The configurations below are historical GPU results from before the tower retirement on July 20, 2026.
| Model | Quant | Gen | Context | Server | Status | Notes |
|---|---|---|---|---|---|---|
| Qwen3.6-27B GPTQ INT4 + MTP n=3 (Genesis) | AutoRound INT4 | ~97 t/s warm (2026-07-08) | 65K | vLLM (Genesis-patched) | production until tower retirement 2026-07-20 | Final local production backend — restored 2026-07-03 after the Ornith trial. GPTQ-Marlin, TP=2, MTP n=3 speculative decoding, chunked prefill + prefix caching. |
| Ornith-1.0-35B | Q4_K_M GGUF / AEON NVFP4 | ~101 t/s warm / 124 peak (prod config) | 131K (NVFP4) | llama.cpp / vLLM 0.23 | trial ended 2026-07-03 | Production trial 2026-06-26 → 2026-07-03. Lab peak 129.9 t/s short-context (2026-06-25). Won sequential tool composition vs DeepSeek V3.2 (20/20 vs 11/20, ~10× faster); DeepSeek won the suite overall by one point. |
| Nemotron 3 Nano 30B (MoE hybrid) | Q4_K_M | 117.34 t/s med / 117.60 peak | 32K | llama.cpp cuda128-clean | historical — measured 2026-05-20 | Mamba/SSM hybrid MoE. ~3B active params/token. Requires cuda128-clean build; cuda13 crashes with MMQ CUDA error on SM_120. |
| AEON (Qwen3 NVFP4) | NVFP4 (ModelOpt) | ~69 t/s | 122K | vLLM 0.9+ | stopped | Blackwell-native quantization. MTP speculative decoding n=3. |
| Qwen3.6-27B (SSM hybrid) | Q4_K_M | ~22 t/s | 65K | llama.cpp | stopped | Mamba/SSM hybrid. Fast prefill (960 t/s), slow gen (SSM bottleneck). |
| Gemma 4 26B (CPU baseline) | Q4_K_M | ~11.7 t/s | 32K | llama.cpp | historical | CPU-only orchestration-tier baseline from the pre-GPU phase. Consistent swap usage 4–8 GB. |
Inference Optimization — Autoresearch Log
130+ automated autoresearch runs across six model architectures using a scripted loop — a subset of the 215+-run total benchmark record. Stopping criterion: +5 tok/s improvement per iteration. Full logs at inference-research.
Key Configurations — Qwen3.6-35B-A3B MoE (dual GPU, April 2026 campaign)
| Configuration | Gen Speed | Prompt Speed | Delta vs Baseline | Notes |
|---|---|---|---|---|
| Dual GPU, all-on-GPU, f16 KV | 107 t/s | 2,436 t/s | +50% vs single GPU | Final config. TP=2, PCIe x8+x4. |
| Single GPU, all-on-GPU, f16 KV | 71 t/s | — | Pre-dual baseline | After CPU offload fix. Clean VRAM fit. |
| Full GPU offload (Exp 2 breakthrough) | 70 t/s | 222 t/s | +118% gen, +200% prompt | vs expert-on-CPU config. |
| Expert tensors on CPU (MoE routing) | 32 t/s | 74 t/s | −55% | CPU bottleneck on MoE expert routing. |
CPU Inference Baseline — Orchestration Tier (17-day log)
| Metric | Value |
|---|---|
| Sample period | April 2–18, 2026 |
| Model | Gemma 4 26B Q4_K_M |
| Avg gen throughput | 10.49 tok/s |
| Range | 5.02 – 11.72 tok/s |
| Avg time to first token | 616 – 2,224 ms |
| Avg swap used | 4.6 – 7.7 GiB |
| Speedup vs GPU tier | 6.6× (10.5 → 74 tok/s) |
Autoresearch Results by Category
Agent Infrastructure
Core agent processes run as hardened background services. Auto-restart, post-boot recovery, and gateway-based inference routing are all part of the runtime design. Slack Socket Mode remains the primary interactive path for the agent stack.
| Process | Framework | Interface | Model tier | Status |
|---|---|---|---|---|
| Agent Harness | Python / custom | Slack Socket Mode | Internal inference gateway | live |
| Mike | RelayV3 (Python) | Discord + Telegram + Slack + IRC | Internal inference gateway | live |
| Harness Personas | Internal harness roles | Not separately exposed | Internal gateway / task-dependent | internal |
| Hermes | Python / hermes-gateway.service | Gateway agent — inference via internal :8010 proxy | Gateway (GPT-5.6 Luna) | live |
| Chronicle | Python / cron | 2am daily cron | Claude API | live |
Service Hardening Details
Raw Data & Artifacts
All benchmark data and logs referenced in the research paper are available publicly.
| Artifact | Location | Contents |
|---|---|---|
| Live benchmark data | boundarylabs.org/benchmarks.html | Inference metrics, MESA scores, LongMemEval results |
| Raw experiment logs | inference-research | 215+ raw benchmark runs, drivers, and verdict docs per research program |
| Agent infrastructure | randomchaos7800-hub | Agent harness, benchmark tooling, selected repos |
| Research paper | papers.html | Full preprint text, tables, references |
| Weekly build logs | dinovitale.com | Weekly recap posts with build decisions and operational notes |