🧬 Hermes Polyglot Interconnect

Models Γ— harnesses Γ— MCP mesh Β· who-calls-whom across your AI estate Β· evidence: session_model_usage (state.db), config.yaml, Cursor pack sha e4886fc2…

β‘  The interconnect mesh β€” who calls whom

LAYER 1 Β· HARNESSSES (the callers) πŸ›°οΈ Hermes Agent Mac desktop Β· cron Β· swarms Β· skills v0.20.5 Β· 206 sessions ⚑ Cursor Cloud Agents Linux VM Β· agentic-moat repo Β· 37 runs ATLAS/Engine-B discipline πŸ€– Claude Code (pending audit) ~/.claude projects Β· mine-local.sh ready 12-month lookback queued πŸ‘€ Commander (human-in-loop) AUTH gates Β· HALT overrides Β· counsel final authority on all lanes LAYER 2 Β· MODELS (via OpenRouter unless noted) GLM-5.2 (z-ai) MAIN brain β€” Hermes default 5,142 calls Β· paid Qwen3-coder-30b DELEGATION subagent brain 247 calls Β· free tier MOA: Nemotron Ultra/Super :free mixture-of-agents aggregator + refs 1,294 + 547 calls stealth/ox-alpha this auditor persona 94 calls DeepSeek v4 flash/pro burst capacity lane 481 calls Gemini 3 Flash Reflector Worker key (direct) 7 calls Β· reflect lane LAYER 3 Β· MCP SERVERS (tool hands) GDELT Cloud (Hermes) Cloudflare-docs (ready) CF bindings/builds/obs InfraDOME FastMCP+Kuzu nighthawk-opendata forge-api / paybridge-api DOME well-known disc. runtime-mcp (NXDOMAIN!) cursor-cloud + subs Walrus/MemWal (mock) needsAuth: 3 CF servers Β· HALT: ntwx-mcp-1.1 deploy AUTH gates everything ⚠ 112 request-dumps Jul–Aug: 32 rate-limit Β· 43 client-error exhaustions Solid = primary route Β· dashed = secondary/conditional All model calls ride OpenRouter except Gemini (Worker-direct) and InfraDOME FastMCP (local stdio) Free-tier dependency: delegation+MOA = 2,088 calls riding throttled :free lanes
Main-model routes Delegation/subagent MOA + AUTH control Cursor→CF MCP Secondary tool routes

β‘‘ Polyglot explorer β€” filter the mesh

β‘’ Model usage ledger β€” actual calls from Hermes state.db

ModelAPI callsRole in polyglotLaneHealth
z-ai/glm-5.25,142Main conversational + task brainOpenRouter paidHEALTHY
nvidia/nemotron-3-ultra-550b:free1,294MOA aggregator (+super 547 refs)OpenRouter freeTHROTTLE RISK
qwen/qwen3-coder-30b-a3b-instruct247delegate_task subagents (16/16 ok)OpenRouter freeTHROTTLE RISK
deepseek v4 flash+pro481Burst overflow laneOpenRouterOK
stealth/ox-alpha94Audit persona (this dashboard's author)OpenRouterOK
google/gemini-3-flash-preview7Reflector reflect-lane via WorkerGemini key directOK

β‘£ Named call routes

#RoutePathStatus
R1Hermes chat β†’ main brainHermes β†’ OpenRouter β†’ GLM-5.2LIVE Β· 5.1k calls
R2Hermes β†’ background swarmsHermes delegate_task β†’ Qwen3-coderLIVE Β· free-tier 429 risk
R3Hermes β†’ MOA synthesisHermes β†’ Nemotron Ultra (aggregator) + GLM/DeepSeek (refs)CONFIGURED Β· free-tier risk
R4Reflector reflect-laneReflector Worker β†’ Gemini 3 Flash (own key)LIVE
R5Cursor Agent β†’ infra toolsCloud Agent β†’ CF docs/bindings/builds/observability MCPsPARTIAL β€” 3 of 4 needsAuth
R6Cursor Agent β†’ graph brainInfraDOME β†’ FastMCP(Kuzu) stdioEXPERIMENTAL Β· prod pending AUTH
R7Agent discovery β†’ DOMEany agent β†’ dome/.well-known/mcp/servers.jsonLIVE Β· advertised
R8PayBridge tool planemerchants β†’ paybridge-api/mcp404 Β· flag OFF
R9NTWX runtime transportagents β†’ runtime-mcp.ml-nightworx.ioNXDOMAIN Β· deploy HALT (no CF token)
R10Hermes intel feedHermes β†’ GDELT Cloud MCP β†’ briefingsREADY but cron paused upstream
Reading the mesh: Layer 1 = who acts (harnesses). Layer 2 = what thinks (models β€” every call metered). Layer 3 = what they touch (MCP tool servers). The single biggest structural risk is that your delegation + MOA capacity rides free-tier lanes (R2/R3) while your main brain is paid β€” one provider throttle away from silent swarm degradation (already seen: 112 request dumps). Second: R8/R9 are advertised transports that don't answer β€” same dead-glass pattern as Pillar B, but for agents.
Data: Hermes session_model_usage + config.yaml βˆͺ Cursor inventory.json (sha e4886fc2…) Β· generated 2026-08-22.