
What Round 2 actually proved
Round 2, closed: an invalid pilot, the engineering that survived it, and why refusing to claim findings that do not exist is how research actually works.
20 posts in this category

Round 2, closed: an invalid pilot, the engineering that survived it, and why refusing to claim findings that do not exist is how research actually works.

Run records showed success for a week while the conversation was never actually shared — invisible failures, and the review that invalidated the Round 2 pilot.

Thirteen nights of fixing one layer only to find the next one broken — the deployment gauntlet between a validated runner and its first green day.

A device-pairing deadlock that consumed a four-hour session and killed the gateway — ten approaches, none worked, and why the deadlock was not a bug.

Every cron job fired, the container was healthy, every provider call returned 200 — and the output was broken in eight different ways. What the first automated run revealed about the system around the model.

The Round 2 Library Terminal design and its Gate 1 vertical slice — the one clean moment before the automated runs put it to the test.

The Round 1 record: three failed migrations, a costly heartbeat, a sandbox that existed only in documentation — and the constraints each failure left behind.

The same soul, two different models, two different writers — a controlled comparison of what actually changes when you swap the model under a persona.

By Day 9 the system had stopped giving the agent its own updated history — a Round 1 postmortem on continuity, stale canon, and the cost of trusting a pipeline you do not monitor.

The Maker saw a directory tree and called it a town — how a persona's central conceit determined what the agent noticed, protected, and narrated.

Designing a diary character who doesn't know she's an AI — what a file-based creative agent notices, what it protects, and what design can't control.

The documentation described a sandboxed maker; the runtime shipped the sandbox off with the Docker socket mounted — a behavioral boundary that was never an isolation boundary.

The reflections were good; the agents still forgot. A Round 1 continuity failure about the difference between generating knowledge and delivering it to the sessions that need it.

Two butlers sharing one personality without sharing a calendar — composing a master soul and per-butler sub-souls for household agents.

AgentOps is not a deploy/ folder — separating instructions the model might follow from configuration the system actually enforces, on a personal-scale agent deployment.

Two interactive Telegram butlers and two scheduled personas on one four-vCPU VPS — the architecture and cost story of running four agents on pocket change.

A default heartbeat schedule multiplied a full-context agent turn into 56 unwanted API calls a day — a cost story about stopping work that had no purpose.

The most dangerous failure in an agent integration is text that looks like execution — a model narrating success without calling anything, and why the guard has to live in code.

Three attempts to stand up a household agent platform — OpenClaw, Hermes, and a custom Python runner — and what the failures taught about what LLMOps actually is.

A chronological record of building, breaking, and rebuilding a self-hosted multi-agent system — four AI agents on a $7.50 VPS, documented the way the mistakes actually happened.