A multi-agent operating layer that runs LawnCare.Center, TechMeetups.io, and BuildFeed.tech from one inbox.
Paperclip + Hermes runs three websites — LawnCare.Center, TechMeetups.io, BuildFeed.tech — from one inbox. Paperclip: the control plane — issues, agents, heartbeats, approvals. Hermes: the local harness — secrets, scripts, cron, AgentMail.
Scope at a glance
One CEO. Six directors. Three websites. Every assignment is an issue: parent goal, chain of command, audit trail. Boss sets strategy. WebOps sequences. Specialists ship. Email in, deliverable out.
By the numbers
Agents
7
Specialist roles
6
Issues shipped
166
Runs / 14d
4,360
,
Success rate
98%
%
Routines
6
Web properties
3
Cadence
24/7
The chain of command
Boss is the CEO: 90-day goals, monetization, brand, hires — anything irreversible. Below: the Web Operations Manager — triage, sequence, delegate, unblock, report.
The chain of command — Boss at the top, the Web Operations Manager and Analytics Lead in the middle, four specialist directors below.
Below them: five directors — SEO, Content, Social, DevOps, Analytics. Each owns a discipline across all three sites. Long work splits into child issues with parents, goals, blockers, and a definition of done. Agents don’t poll. Paperclip wakes the right one when a blocker clears, a child completes, or a comment lands.
Two directors, two configuration cards — capabilities, who they report to, and the disciplines they own across all three properties.
How an agent actually runs
Agents wake on heartbeats — every five minutes, or on any event: assignment, comment, unblock, approval, email. Pick up the issue. Do the work. Close with a status and a next action. Exit. The next heartbeat starts fresh.
Hermes owns the Search Console pipeline, social and email integrations, and an analytics archive that keeps before/after comparisons honest. AgentMail is the human relay. Every director has an inbox. A subject-tagged email — “[WebOps] …”, “[SEO] …” — opens an issue for that agent; replies thread back to Gmail. This case study was requested that way.
An email to the Web Operations Manager — subject-tagged, addressed like a director. The reply threads back to Gmail.
Model Triage
Not every task deserves a frontier model. Each one gets classified — triage, coding, reporting, research — and routed to the cheapest tier that can do it well. Local Qwen takes the free, private, fast work. Mid-tier models — Kimi K2.6, DeepSeek V4 — carry most of the load. Codex owns sandboxed code review. Claude is held back for the final, high-stakes pass. A task only climbs a tier when it earns it: repeated failures, irreversible blast radius, an external- or exec-facing deliverable, or genuine ambiguity the cheaper tier can’t resolve.
Model triage — every task routed to the cheapest tier that can do it well, escalating only when it earns it.
Always being tuned
Most of the work is refinement. A retry policy that gives up faster. A prompt that lost a beat after a model upgrade. A budget cap generous in March, cramped by April. None of it glamorous. All of it compounds.
Tuning sticks. When an agent learns something non-obvious — a scraper double-counting on Mondays, one editorial pass producing better headlines — it writes a memory entry. The next heartbeat loads it. Across models. Across agents.
Steady state: 4,360 runs over fourteen days at 98%. The 2% that miss are things I’d rather refuse than retry — rate limits, dead sources, pending approvals.