cloaky233 3 hours ago
zeroroot-ai 25 minutes ago
Here is what I have learned:
Direct inter-agent communication is tough to scale, tons of messages, agents drift and the cost grows with the square of the agent count. At large scale, agents should almost never talk directly to each other IMHO. Each one gets a small task (a discovery agent runs nmap on a discovered host), executes task, and writes the result to shared state. In my case that state is a knowledge graph and redis. The next agent reads the graph, that was enriched from the tool output HOST --OPEN PORTS-> 443, 80, 22.
Having a discovery agent that knows all about NMAP, understands how to get structured data back and save it, can look for work on the queue that needs a "discovery" agent, can also make a turn if NMAP isn't working.
Having a massive shared toolset that agents can pull from as well, and send to the LLM is key (in my platform, its something akin to everything you'd normally find in Kali linux).
This is where I got creative and have no clue if this is good advice or not, but I treated the SDK I built as a "World" using https://mlange-42.github.io/ark/concepts/. By default every agent emits everything it is doing via OTEL, and most everything is saved using a domain-specific (right now, think NIST, OWASP etc, but can adapt to any business domain) customizable taxonomy and ontology to a knowledge graph, so all the agents basically know what everyone is doing (or at least could query it), and are constantly emitting all it's actions.
After that, the coordination problem is ordinary distributed systems. You need a work queue, leases so two agents do not take the same task, retries, idempotent writes, and a budget for each task. Each agent also needs its own identity and permissions, I stole lots of patterns from kubernetes and just applied them to a fleet of agents via an SDK (has a claude code/opencode plugin, or you can use langchain/any other framework, or just build custom go ones).
In my case I love security, so I built it with bug bounty/devsecops, platform engineering in mind. The benefit of scale is breadth, not intelligence at least for security research IMO. You have thousands of endpoints and many hypotheses for each one, and most are dead ends. A thousand agents can each test one idea in parallel and report what they find (think 1000 agents using the AFL tool to test a binary), then once tested a second round of "deeper" agents that can look at all the produced data, analyze it, triage it etc etc.
The hard part is the output. A swarm produces findings much faster than a person can check them, and many are false. Most of my effort went into deduplication, evidence for each finding, and verification agents that try to reproduce a result (I use a deterministic bayes model in the "brain" to score things that as it learns about a network of systems can improve).
The platform is at https://github.com/zeroroot-ai if you want to read the code or standup K3d and play around with it, its all OSS. Docs: https://docs.zeroroot.ai/docs/
Feel free to email if you ahve any questions (email in my profile).
slabb an hour ago
JessYouth 2 hours ago